# New merged nodes from existing nodes

**URL:** <https://community.neo4j.com/t/new-merged-nodes-from-existing-nodes/72783>\
**Category:** Newbie Questions\
**Tags:** apoc, cypher\
**Created:** [February 24, 2025, 6:24pm UTC](https://community.neo4j.com/t/new-merged-nodes-from-existing-nodes/72783 "2025-02-24T18:24:40Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![bologna3](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/bologna3/32/34142_2.png) [@bologna3](https://community.neo4j.com/u/bologna3)\
**Post date:** [February 24, 2025, 6:24pm UTC](https://community.neo4j.com/t/new-merged-nodes-from-existing-nodes/72783/1 "2025-02-24T18:24:41Z")

</div>

Hi all,

Neo4j beginner here. This seems like it should be a simple task, but I really cannot figure it out. I want to create a new set of nodes based off an existing set, merging them based on a unique value for a certain variable.

I have 3.7 million nodes of type "person" successfully loaded into my database. They have several variables attached. I want to create a new set of nodes called "name," where there is one node per unique value of the variable "person\_name\_cluster\_key" in the "person" nodes. (Based on working with the data in R, I know that this should result in 2.3 million "name" nodes). I also want to bring over another variable called "person\_name" for each of the new "name" nodes. Of the multiple "person" nodes merged, I don't care which "person" node this value is take from. Then, I want to relate each new "name" node to the original "person" nodes with a (n:name)-[:name\_of]-\>(p:person) relationship.

I need the process to be iterative and computationally efficient since the dataset is so large. I feel like this should be really simple, but I'm stumped.

Thanks so much.

---

<div class="post-metadata">

**Author:** ![glilienfield](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/glilienfield/32/27534_2.png) [@glilienfield](https://community.neo4j.com/u/glilienfield)\
**Post date:** [February 25, 2025, 2:11am UTC](https://community.neo4j.com/t/new-merged-nodes-from-existing-nodes/72783/2 "2025-02-25T02:11:08Z")

</div>

You can try something like this. Let's see how it works.

Test data:

```auto
unwind [["santa","a"], ["pluto","b"], ["goofy","c"], ["rudolph","d"], ["micky","e"], ["rudolph","f"], ["pluto","g"], ["minnie","h"], ["micky","i"]] as person
create (:Person{person_name_cluster_key: person[0], person_name: person[1]})

```

 ![Screenshot 2025-02-24 at 9.04.45 PM](https://us1.discourse-cdn.com/flex021/uploads/neo4jcommunity/original/3X/f/3/f32fe33204ff8002d2e4caf610ca7a5757344e25.png)

Query:

```auto
:auto
CALL () {
    MATCH(p:Person) 
    WITH p.person_name_cluster_key as key, collect(p) as persons_for_key
    CREATE (n:Name{person_name_cluster_key: key, person_name: head(persons_for_key).person_name})
    FOREACH(i in persons_for_key |
        MERGE (n)-[:name_of]->(i) 
    )
} in CONCURRENT TRANSACTIONS of 100000 ROWS

```

 ![Screenshot 2025-02-24 at 9.07.23 PM](https://us1.discourse-cdn.com/flex021/uploads/neo4jcommunity/original/3X/2/d/2d7a77fa94a5a4a166a54f78c655a1b140b4756a.jpeg)

Note: 1 you need the ":auto" when executing the query in the browser or cypher-shell, but not otherwise.

Note 2: You definitely need an indexes on the labels and properties you are matching and merging on. This is not an issue with this query, as you are not matching on a specify property.

Note 3: I have not used the CONCURRENT TRANSACTIONS clause. It is relatively new.

Note 4: This may require a lot of memory due to the size of your database and the use of a COLLECT on your entire data set.

Let me know how it goes.
