# Delete duplicate data and restore relationship

**URL:** <https://community.neo4j.com/t/delete-duplicate-data-and-restore-relationship/15950>\
**Category:** Cypher\
**Tags:** cypher\
**Created:** [March 17, 2020, 1:51pm UTC](https://community.neo4j.com/t/delete-duplicate-data-and-restore-relationship/15950 "2020-03-17T13:51:21Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![benjamin.caoduro\_ext](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/benjamin.caoduro_ext/32/3922_2.png) [@benjamin.caoduro\_ext](https://community.neo4j.com/u/benjamin.caoduro_ext)\
**Post date:** [March 17, 2020, 1:51pm UTC](https://community.neo4j.com/t/delete-duplicate-data-and-restore-relationship/15950/1 "2020-03-17T13:51:21Z")

</div>

Hi everyone !

I have a problem and i would like to know if there is a sample way to solve it.

I have a function under java which create nodes with several properties, then i would like to compare those created nodes (their labels, their properties and content (except one property : the Id property which is unique)) with nodes in my data base using a cypher command. It needs to be dynamic, I mean the properties and labels are different according to the data created (but they all have an Id property).

If I found duplicate nodes (except Id property and the content) I would like to delete duplicate nodes, but before deleting them I would like to add the relationship they have to the unique node which I will not delete.

---

<div class="post-metadata">

**Author:** ![elena.kohlwey](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/elena.kohlwey/32/31033_2.png) [@elena.kohlwey](https://community.neo4j.com/u/elena.kohlwey)\
**Post date:** [March 17, 2020, 3:54pm UTC](https://community.neo4j.com/t/delete-duplicate-data-and-restore-relationship/15950/2 "2020-03-17T15:54:00Z")

</div>

Hi Benjamin,

for your first question: have you looked into node similarity algorithms yet? ([Node Similarity - Neo4j Graph Data Science](https://neo4j.com/docs/graph-algorithms/current/algorithms/node-similarity/))

for the second: if you create a relationship from a duplicate node to the original node and then delete the duplicate node, then the relationship will disappear as well... a relationship always needs two nodes to cling to...

So, if we assume that there is no relationship that you can keep from the duplicate node, might a solution be to use the "MERGE" command and not the "CREATE" command in your java programme? This will then only create a node if there is no node yet...  
For your "keeping track" (which I guess is what you wanted to do with the relationship) would it be a solution to "MATCH" for the node and see if it already exists and then add some kind of property to it?

I hope my thoughts help you.  
Regards,  
Elena

---

<div class="post-metadata">

**Author:** ![intouch\_vivek](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/intouch_vivek/32/4097_2.png) [@intouch\_vivek](https://community.neo4j.com/u/intouch_vivek)\
**Post date:** [March 17, 2020, 4:33pm UTC](https://community.neo4j.com/t/delete-duplicate-data-and-restore-relationship/15950/3 "2020-03-17T16:33:28Z")

</div>

While creating a node try Merge instead of Create.  
Merge (a:Node1{id:line.node\_id}) On Create set a.name =line. name, a.startdate=line.startDate  
Return a

Else If you have already nodes in the Graph then use try to merge the nodes  
MATCH (a:Node1)  
WITH a.id as id, collect(a) as nodes  
CALL apoc.refactor.mergeNodes(nodes, {properties: "combine"}) YIELD node  
RETURN node;
