Duplicate Record Cleanup in Production MongoDB by Kellyn LiDuplicate Record Cleanup in Production MongoDB by Kellyn Li

Duplicate Record Cleanup in Production MongoDB

Kellyn Li

Kellyn Li

Safely Cleaning Duplicate Production Data

Resolving 100+ duplicate records across interconnected MongoDB collections without breaking a live SaaS system.

The situation

While working on a parish-management feature, I discovered that 100+ duplicate parish records had been imported into production and had already been in use for several months.
The problem was larger than simply deleting duplicates. These records were referenced across multiple MongoDB collections, including a core Person collection whose backup alone exceeded 90 GB. Removing the wrong record could leave references inconsistent across the system. Projects

My role

I investigated the scope of the data issue and designed the cleanup strategy.
The objective was to restore data consistency while minimizing how much existing production data needed to be changed.

Investigation

Before modifying anything, I first needed to understand the blast radius.
I scanned the related collections to determine where the duplicate records were being referenced and how extensively each duplicate had already been used.
This turned the problem from:
“Which duplicates should we delete?”
into:
“Which record should remain so that we disturb the least amount of existing production data?”

What I found

The duplicate records were not equally important.
Some had accumulated significantly more references than others during the months they had been in production. Treating every duplicate equally would therefore create unnecessary remapping across the database.

The strategy

I designed a mechanism to identify the most-used record as the canonical record, preserve it, and consolidate the remaining duplicates around it.
Instead of rewriting everything, the cleanup targeted only the records and references that actually needed to change.
The principle was simple:
Understand the relationships first → preserve the most established record → modify the minimum necessary data.

Outcome

The cleanup scanned 60,000+ documents, while only around 3% needed to be modified, remapped, or archived.
The production operation completed in under 40 minutes. Projects

What this case demonstrates

This wasn't primarily a database-cleanup task. The difficult part was deciding what could safely be changed in an interconnected production system.
The workflow was:
Unexpected data issue → Scope the impact → Map dependencies → Design minimum-impact strategy → Execute cleanup
Role: Software Engineer Focus: Production data · Risk assessment · Data remediation · Problem diagnosis Tools: MongoDB · C# · Production data scripts
Like this project

Posted Sep 24, 2026

Resolved duplicate records in MongoDB for a SaaS system without disrupting production.