# From Python notebook to Neo4j graph via Cypher query

**URL:** <https://community.neo4j.com/t/from-python-notebook-to-neo4j-graph-via-cypher-query/56415>\
**Category:** Cypher\
**Tags:** cypher\
**Created:** [May 16, 2022, 10:13pm UTC](https://community.neo4j.com/t/from-python-notebook-to-neo4j-graph-via-cypher-query/56415 "2022-05-16T22:13:25Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![jatinjaitleypro](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/jatinjaitleypro/32/21860_2.png) [@jatinjaitleypro](https://community.neo4j.com/u/jatinjaitleypro)\
**Post date:** [May 16, 2022, 10:13pm UTC](https://community.neo4j.com/t/from-python-notebook-to-neo4j-graph-via-cypher-query/56415/1 "2022-05-16T22:13:25Z")

</div>

I have a dataframe which I wish to transfer from python to Neo4j. My dataframe looks like below.

![image](https://us1.discourse-cdn.com/flex021/uploads/neo4jcommunity/original/3X/9/6/96791b889a34104272eac29694e7c05f7d7b32f3.png)

I want the text column to be connected via Next relationship. Something like below.

 ![image](https://us1.discourse-cdn.com/flex021/uploads/neo4jcommunity/original/3X/0/8/082559bc2ac505754b7236dad31fd046768b168e.png)

I know the Cypher query. My requirement is I want the POS column rows attached as a property to each word. Example Node Dog has has POS NOUN so NOUN should be attached as a property to that node and the NEXT relationship should be maintained as shown above.

How can I write the query in python notebook and see the same results in Neo4j graph? Please assist me with the syntax as I am pretty new to Neo4j and Cypher?

---

<div class="post-metadata">

**Author:** ![jatinjaitleypro](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/jatinjaitleypro/32/21860_2.png) [@jatinjaitleypro](https://community.neo4j.com/u/jatinjaitleypro)\
**Post date:** [May 18, 2022, 9:07am UTC](https://community.neo4j.com/t/from-python-notebook-to-neo4j-graph-via-cypher-query/56415/2 "2022-05-18T09:07:04Z")

</div>

Here is what I have done so far.  
I have py2neo and neo4j installed in my PC.  
I wish to run the cypher query from my python notebook and the changes should reflect in NEO4j graph  
Already know some basic like

```auto
import pandas as pd
from py2neo import Graph,Node,Relationship
from neo4j import GraphDatabase, basic_auth
graph = Graph("http://localhost:7474/browser/", auth=("neo4j", " *****"))

for index, row in df.iterrows():
    tx = graph.begin()
    tx.evaluate('''cypher query goes here''')
    tx.commit()

```

From python notebook by using a dataframe putting the value of second column POS as property and maintaining the Next relationship in the first column as shown above

---

<div class="post-metadata">

**Author:** ![koji](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/koji/32/27564_2.png) [@koji](https://community.neo4j.com/u/koji)\
**Post date:** [May 18, 2022, 1:08pm UTC](https://community.neo4j.com/t/from-python-notebook-to-neo4j-graph-via-cypher-query/56415/3 "2022-05-18T13:08:36Z")

</div>

Hi @jatinjaitleypro

How about this.  
I just added pos.

```auto
WITH split(tolower("His dog eats turkey on Tuesday")," ") AS text,
     split("PRON NOUN VERB PROPN ADP PROPN"," ") AS pos
UNWIND range(0,size(text)-2) AS i
MERGE (w1:Word {name: text[i], pos: pos[i]})
MERGE (w2:Word {name: text[i+1], pos: pos[i+1]})
MERGE (w1)-[:NEXT]->(w2)
RETURN w1, w2

```

---

<div class="post-metadata">

**Author:** ![jatinjaitleypro](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/jatinjaitleypro/32/21860_2.png) [@jatinjaitleypro](https://community.neo4j.com/u/jatinjaitleypro)\
**Post date:** [May 18, 2022, 1:17pm UTC](https://community.neo4j.com/t/from-python-notebook-to-neo4j-graph-via-cypher-query/56415/4 "2022-05-18T13:17:30Z")

</div>

@koji No, I am not looking for this. The challenge I am facing is on jupyter notebook how do I perform the same operation via pandas dataframe from python notebook

```auto
import pandas as pd
from py2neo import Graph,Node,Relationship
from neo4j import GraphDatabase, basic_auth
graph = Graph("http://localhost:7474/browser/", auth=("neo4j", " *****"))

for index, row in df.iterrows():
    tx = graph.begin()
    tx.evaluate('''cypher query goes here''')
    tx.commit()
```

---

<div class="post-metadata">

**Author:** ![paltusplintus](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/paltusplintus/32/13463_2.png) [@paltusplintus](https://community.neo4j.com/u/paltusplintus)\
**Post date:** [May 23, 2022, 12:35pm UTC](https://community.neo4j.com/t/from-python-notebook-to-neo4j-graph-via-cypher-query/56415/5 "2022-05-23T12:35:36Z")

</div>

[neointerface](https://github.com/GSK-Biostatistics/neointerface) package has load\_df method, however in order to create NEXT relationships btw words you need to load the dataframe with index column and run an additional query. try something like:

```auto
#pip install neointerface
import neointerface
import pandas as pd
db = neointerface.NeoInterface(host="neo4j://localhost:7687" , credentials=("neo4j", "YOUR_NEO4J_PASSWORD"))
df = pd.DataFrame(...)
db.load_df(df.reset_index(), label="Word", merge=False)
db.create_index("Word", "index")
db.query("MATCH (w1:Word), (w2:Word) WHERE w2.index = w1.index + 1 MERGE (w1)-[:NEXT]->(w2)")

```

---

<div class="post-metadata">

**Author:** ![jatinjaitleypro](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/jatinjaitleypro/32/21860_2.png) [@jatinjaitleypro](https://community.neo4j.com/u/jatinjaitleypro)\
**Post date:** [May 29, 2022, 5:37pm UTC](https://community.neo4j.com/t/from-python-notebook-to-neo4j-graph-via-cypher-query/56415/6 "2022-05-29T17:37:28Z")

</div>

> [@paltusplintus](#):
>
> ```auto
> import pandas as pd
> db = neointerface.NeoInterface(host="neo4j://localhost:7687" , credentials=("neo4j", "YOUR_NEO4J_PASSWORD"))
> df = pd.DataFrame(...)
> db.load_df(df.reset_index(), label="Word", merge=False)
> db.create_index("Word", "index")
> db.query("MATCH (w1:Word), (w2:Word) WHERE w2.index = w1.index + 1 MERGE (w1)-[:NEXT]->(w2)")
> 
> ```

@paltusplintus

First of all thank you for your response I was able to use "neointerface".  
But my goal is still not achieved. Here is what I have tried.

 ![image](https://us1.discourse-cdn.com/flex021/uploads/neo4jcommunity/original/3X/c/d/cd4390247a606fd0a95b1ca943e930ade07c0f51.png)

Then

 ![image](https://us1.discourse-cdn.com/flex021/uploads/neo4jcommunity/original/3X/6/f/6f718b7082ce41e09de8d4bbcb3a61f53a1ade62.png)

Output:

 ![image](https://us1.discourse-cdn.com/flex021/uploads/neo4jcommunity/original/3X/d/6/d6fdece86c5dbc36d64c4417ac07570115a12212.png)

Ideally the node the and "The" should have been created once but they were created twice.  
example

 ![image](https://us1.discourse-cdn.com/flex021/uploads/neo4jcommunity/original/3X/8/8/88005acbe7270cb4bb61295f79ad7044de85d456.png)

and

 ![image](https://us1.discourse-cdn.com/flex021/uploads/neo4jcommunity/original/3X/0/9/092fd6b917a10740ad055a489f704210635ac6d9.png)

**What I am looking for is there should be no duplicate text**  
**The POS against each text word should be created as a list or array or collection anything**

Example: "text" : "wild" (should only be one node)  
"POS": ["NOUN", "ADJ"]

---

<div class="post-metadata">

**Author:** ![paltusplintus](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/paltusplintus/32/13463_2.png) [@paltusplintus](https://community.neo4j.com/u/paltusplintus)\
**Post date:** [May 30, 2022, 11:18am UTC](https://community.neo4j.com/t/from-python-notebook-to-neo4j-graph-via-cypher-query/56415/7 "2022-05-30T11:18:49Z")

</div>

The load\_df will not work for you in this specific use case. Solved your problem as follows. Note that you need to install apoc library to make it work:

```auto
import spacy
import en_core_web_sm
import pandas as pd
nlp = spacy.load("en_core_web_sm")

text = "The wild is dangerous"
doc = nlp(text)
cols = ("text", "POS")
rows = []
for t in doc:
    row = [t.text, t.pos_]
    rows.append(row)
df = pd.DataFrame(rows, columns = cols)

#clean-up
db.query("MATCH (w:Word) detach delete w")
db.create_index("Word", "index")

q = """
UNWIND $data as row
MERGE (w:Word{text:row.text})
ON CREATE SET w.POS = [row.POS]
ON MATCH SET w.POS = CASE WHEN row.POS in w.POS THEN w.POS ELSE w.POS + [row.POS] END
WITH collect(w) as coll
WITH apoc.coll.pairsMin(coll) as pairs
UNWIND pairs as pair
WITH pair[0] as node1, pair[1] as node2
MERGE (node1)-[:NEXT]->(node2)
"""

db.query(q, {'data': df.to_dict(orient='records')})

text = "The rockstar is wild"
doc = nlp(text)
cols = ("text", "POS")
rows = []
for t in doc:
    row = [t.text, t.pos_]
    rows.append(row)
df_2 = pd.DataFrame(rows, columns = cols)

db.query(q, {'data': df_2.to_dict(orient='records')})

```
