Original synthetic demonstration. Created from fictional records to demonstrate a controlled cleanup workflow; this is not paid client work.
The problem
A small export contains an exact duplicate and inconsistent spacing in customer names. It also contains details that a careless cleanup could damage: an ID with leading zeros, a note spanning two lines, meaningful spaces inside another note, and blank values.
The approach
I set narrow rules before changing the data:
Remove a record only when the entire original row is an exact duplicate.
Trim outer whitespace only in the customer-name field.
Preserve identifiers as text, including 00017.
Keep multiline notes, note whitespace and blank values intact.
Write a separate cleaned output and record each change.
What the sample demonstrates
The five-row source becomes a four-row cleaned file. The change log records exactly three actions: two customer-name whitespace trims and one exact duplicate omission.
The leading-zero identifier remains unchanged. The two-line note stays on two lines, the padded note keeps its spaces, and blank values remain blank. These facts were checked against the source, output and change log.
The handoff
The demonstration package contains the original CSV, the cleaned CSV, a machine-readable change log and a plain-language explanation of the rules. A buyer can inspect what changed and why.
Applying this to a real export
A real project starts with a representative sample and agreement on the desired output. Duplicate definitions, formatting rules and exceptions are confirmed before processing. Ambiguous values are flagged for a decision instead of guessed.
This example demonstrates the workflow on a small fictional dataset. It does not claim client savings, business results, universal duplicate detection, or large-scale performance.
Like this project
Posted Sep 7, 2026
Original synthetic demonstration: controlled CSV cleanup, a reviewable change log, and checks that preserve IDs, multiline notes and blank values.