# Not detecting repeated nodes

**URL:** <https://community.neo4j.com/t/not-detecting-repeated-nodes/58441>\
**Category:** Neo4j Graph Platform\
**Tags:** migrated\
**Created:** [January 20, 2023, 4:48am UTC](https://community.neo4j.com/t/not-detecting-repeated-nodes/58441 "2023-01-20T04:48:29Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Reuben](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/reuben/32/27511_2.png) [@Reuben](https://community.neo4j.com/u/Reuben)\
**Post date:** [January 20, 2023, 4:48am UTC](https://community.neo4j.com/t/not-detecting-repeated-nodes/58441/1 "2023-01-20T04:48:29Z")

</div>

My cypher query is not able to detect duplicates / repeated nodes under different labels. The output it gives me is –\> (no changes, no records)

```auto
MATCH (n)
WHERE n.name = "Joining"
WITH n, COUNT(n) as count
WHERE count > 1
RETURN n

```

However, when search for the same node as an individual entity using this query below gives me all the available duplicates under the different labels.

```auto
MATCH (n)
WHERE n.name = "Joining"
// WITH n, COUNT(n) as count
// WHERE count > 1
RETURN n

```

Please can anyone explain why and suggest how best to go about it? Thanks

#Neo4J #Cypher #nodeduplicates

[@glilienfield](https://community.neo4j.com/t5/user/viewprofilepage/user-id/9010)

---

<div class="post-metadata">

**Author:** ![Reuben](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/reuben/32/27511_2.png) [@Reuben](https://community.neo4j.com/u/Reuben)\
**Post date:** [January 20, 2023, 6:18am UTC](https://community.neo4j.com/t/not-detecting-repeated-nodes/58441/2 "2023-01-20T06:18:09Z")

</div>

Thank you as usual [@glilienfield](https://community.neo4j.com/t5/user/viewprofilepage/user-id/9010)

---

<div class="post-metadata">

**Author:** ![Reuben](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/reuben/32/27511_2.png) [@Reuben](https://community.neo4j.com/u/Reuben)\
**Post date:** [January 20, 2023, 5:15am UTC](https://community.neo4j.com/t/not-detecting-repeated-nodes/58441/3 "2023-01-20T05:15:59Z")

</div>

Is there a way to look for duplicate nodes in general? Something like this?

```auto
// this doesn't work though*
MATCH (n)
WITH n, COUNT(n) as count
WHERE count > 1
WITH COLLECT(n) as nodes
RETURN nodes

```

---

<div class="post-metadata">

**Author:** ![Reuben](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/reuben/32/27511_2.png) [@Reuben](https://community.neo4j.com/u/Reuben)\
**Post date:** [January 20, 2023, 5:00am UTC](https://community.neo4j.com/t/not-detecting-repeated-nodes/58441/4 "2023-01-20T05:00:40Z")

</div>

using the collect approach worked:

```auto
MATCH (n)
WHERE n.name = "specific name"
WITH COLLECT(n) as nodes
WHERE SIZE(nodes) > 1
RETURN nodes

```

---

<div class="post-metadata">

**Author:** ![glilienfield](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/glilienfield/32/27534_2.png) [@glilienfield](https://community.neo4j.com/u/glilienfield)\
**Post date:** [January 20, 2023, 4:57am UTC](https://community.neo4j.com/t/not-detecting-repeated-nodes/58441/5 "2023-01-20T04:57:56Z")

</div>

Simple mistake, you are grouping by 'n', which is the node. The result will be a separate row for each n and a corresponding count of one.

What you want to do is group on the common value, which is n.name.

```auto
MATCH (n)
WHERE n.name = "Joining"
WITH n.name as name, COLLECT(n) as duplicates
WHERE size(duplicates) > 1
RETURN duplicates

```

---

<div class="post-metadata">

**Author:** ![glilienfield](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/glilienfield/32/27534_2.png) [@glilienfield](https://community.neo4j.com/u/glilienfield)\
**Post date:** [January 20, 2023, 6:02am UTC](https://community.neo4j.com/t/not-detecting-repeated-nodes/58441/6 "2023-01-20T06:02:19Z")

</div>

You need to group the nodes by their duplicate values. Assume that you define two nodes as duplicates if their duplicateProperty1 and duplicateProperty2 values are equal. For example, two Person nodes are duplicate if they have the same email address and phone numbers, something like that. Using properties duplicateProperty1 and duplicateProperty2 as the values that define a duplicate, the following query will group the nodes that have the same values and return their node ids in a list. You can return what you need instead.

```auto
match(n)
with n.duplicateProperty1 as property1, n.duplicateProperty2 as property2, collect(id(n)) as ids
where size(ids) > 1
return property1, property2, ids

```

Test data:

![Screen Shot 2023-01-20 at 12.58.26 AM.png](https://us1.discourse-cdn.com/flex021/uploads/neo4jcommunity/original/3X/8/2/8276996cd54d05a858f52fc4cef8f0f7ce185d8e.png "Screen Shot 2023-01-20 at 12.58.26 AM.png")

Result of grouping by both properties:

![Screen Shot 2023-01-20 at 12.58.51 AM.png](https://us1.discourse-cdn.com/flex021/uploads/neo4jcommunity/original/3X/4/9/4914cb1cba35690e4dc48c6e7a43200efd79167e.png "Screen Shot 2023-01-20 at 12.58.51 AM.png")

---

<div class="post-metadata">

**Author:** ![glilienfield](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/glilienfield/32/27534_2.png) [@glilienfield](https://community.neo4j.com/u/glilienfield)\
**Post date:** [January 20, 2023, 5:02am UTC](https://community.neo4j.com/t/not-detecting-repeated-nodes/58441/7 "2023-01-20T05:02:57Z")

</div>

You know you are correct, the 'with c.name' as I did is not necessary since the value is the same for all nodes. As such, it can be removed, which is the result you gave.

---

<div class="post-metadata">

**Author:** ![Reuben](https://sea1.discourse-cdn.com/flex021/user_avatar/community.neo4j.com/reuben/32/27511_2.png) [@Reuben](https://community.neo4j.com/u/Reuben)\
**Post date:** [January 20, 2023, 5:34am UTC](https://community.neo4j.com/t/not-detecting-repeated-nodes/58441/8 "2023-01-20T05:34:32Z")

</div>

Thanks so much, you are a life saver. how about a general scenario to check duplicates as I illustrated below?
