Audit one CSV on your own computer with a repeatable Python command.
This utility removes exact field-for-field duplicates and optionally flags conflicting or blank IDs for review. Parsed cell strings stay unchanged, including leading zeros and numeric-looking text. Uncertain records are kept visible rather than guessed or corrected.
Each run creates cleaned.csv, duplicates.csv, review.csv, audit.json and a byte-for-byte copy of the original input. The audit reconciles record counts and records source references and file hashes.
Included: Python source, English run instructions, five executable CLI tests, and a synthetic 12-column demonstration with its output files. The demo has 12 records: 5 cleaned, 5 for review and 2 exact duplicates.
Requires Python 3.10+ and use of a terminal. Input must be one UTF-8 comma-delimited CSV with unique, nonempty headers. No extra Python packages, API key or network connection is required.
Verified locally on macOS, including 10,000 records and a 14-column schema. Windows and Linux have not been tested natively for this release. Configurable safety caps are not performance guarantees.
Scope: exact duplicates and optional ID review, one file at a time. The utility does not infer business values, repair ambiguous records, match approximate duplicates, merge files, or create and refresh Excel workbooks.
Use permission: The purchaser may run and modify this copy for their own personal or internal business work. Client-facing data-processing services using the tool are allowed. Redistribution, resale of the package or modified copies, sublicensing, and uploading the source to a public repository are not included.
The checkout information and included LICENSE.txt describe the 14-day fault support and refund-review conditions.