# Performance Issues Merging Nodes

**URL:** <https://community.neo4j.com/t/performance-issues-merging-nodes/53082>\
**Category:** Cypher\
**Tags:** apoc, performance, cypher\
**Created:** [March 8, 2022, 12:18pm UTC](https://community.neo4j.com/t/performance-issues-merging-nodes/53082 "2022-03-08T12:18:01Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![rezaahnadi99887](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/rezaahnadi99887/32/22847_2.png) [@rezaahnadi99887](https://community.neo4j.com/u/rezaahnadi99887)\
**Post date:** [March 8, 2022, 12:18pm UTC](https://community.neo4j.com/t/performance-issues-merging-nodes/53082/1 "2022-03-08T12:18:01Z")

</div>

Hi,  
I have a database with 20 million nodes and 10 million relationships. I want to merge nodes that have the same code number property.

My cypher is like this

```auto

CALL apoc.periodic.iterate("
MATCH (n:Person) with distinct n.code as props return props
","
UNWIND props as prop
CALL{
WITH prop
MATCH (n:Person {code:prop}) 
with COLLECT(n) AS ns, count(n) as cn where cn > 1
CALL apoc.refactor.mergeNodes(ns, {properties:{OtherCodes:'combine', `.*`: 'overwrite'}}) 
YIELD node RETURN node as s
}WITH s
RETURN s;
", {batchSize:10000, parallel:true, iterateList:true});

```

But it does nothing, does not exist any errors, but it does not process

I use the 4.3.6 neo4j version

---

<div class="post-metadata">

**Author:** ![michael.hunger](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/michael.hunger/32/27377_2.png) [@michael.hunger](https://community.neo4j.com/u/michael.hunger)\
**Post date:** [March 10, 2022, 8:33pm UTC](https://community.neo4j.com/t/performance-issues-merging-nodes/53082/2 "2022-03-10T20:33:37Z")

</div>

Remove the unwind from your 2nd statement.

I presume you have an index on :Person(code) ?

You don't need the subquery.

How many people with the same code do you have 10, 100, 10000 ?

````auto
CALL apoc.periodic.iterate("
MATCH (n:Person) return distinct n.code as prop
","
MATCH (n:Person {code:prop}) 
with prop, COLLECT(n) AS ns, count(n) as cn where cn > 1
CALL apoc.refactor.mergeNodes(ns, {properties:{OtherCodes:'combine', `.*`: 'overwrite'}}) 
YIELD node 
RETURN count(*)
", {batchSize:100, parallel:true, iterateList:true});```
````

---

<div class="post-metadata">

**Author:** ![rezaahnadi99887](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/rezaahnadi99887/32/22847_2.png) [@rezaahnadi99887](https://community.neo4j.com/u/rezaahnadi99887)\
**Post date:** [March 13, 2022, 12:31pm UTC](https://community.neo4j.com/t/performance-issues-merging-nodes/53082/3 "2022-03-13T12:31:48Z")

</div>

Thanks for your answer  
yes I have an index on Person(code)  
And the number of people with the same code is about 100,000

---

<div class="post-metadata">

**Author:** ![rezaahnadi99887](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/rezaahnadi99887/32/22847_2.png) [@rezaahnadi99887](https://community.neo4j.com/u/rezaahnadi99887)\
**Post date:** [March 13, 2022, 12:48pm UTC](https://community.neo4j.com/t/performance-issues-merging-nodes/53082/4 "2022-03-13T12:48:31Z")

</div>

I have a similar problem with relationships.  
Because of the error "All Relationships must have the same start and end nodes.", I wrote a function in my plugin to categorize relationships by start and end, and it works fine.

My cypher is like this

```auto
CALL apoc.periodic.iterate("
MATCH (s:Person)-[r:Work]-(t:Office) WITH COLLECT(r) as lrs RETURN lrs
","
with customPlugin.relations.groupByStartAndEnd(lrs) as grs
UNWIND grs as gr
CALL apoc.refactor.mergeRelationships(gr) YIELD rel RETURN rel
", {batchSize:500, parallel:true, iterateList:true});

```

But when use `apoc.refactor.mergeRelationships` after a while, I see the following error in the logs

_Exception: java.lang.OutOfMemoryError thrown from the UncaughtExceptionHandler in thread "neo4j.Scheduler-1"_
