Self-directed technical demonstration using synthetic data; not a paid client engagement.
Starting state
180 supplier-price rows with duplicate keys, inconsistent whitespace and case, mixed country, currency and date formats, invalid quantities, and missing lead times.
Transformation
Raw input was preserved. Deterministic rules normalized fields, composite keys identified duplicates, invalid values entered an exception queue, and every removal or change remained traceable.
Delivered state
150 standardized, import-ready rows; 30 duplicate removals with an audit trail; five editable sheets covering raw data, rules, cleaned output, QA actions, and reconciliation.
Quality boundary
Unverifiable values were never guessed. They remained explicitly unresolved, so a human or downstream system can distinguish evidence, transformation, and uncertainty.
The principle is simple: a dataset is not clean merely because it looks uniform. It is clean when material changes can be inspected and reproduced.
Self-directed case using synthetic data: 180 messy rows became 150 import-ready rows, with raw evidence, deterministic rules, QA logs and unresolved values.