sample_data/messy_sales.csv. It has duplicate rows, missing values, inconsistent casing and one obvious outlier built in.cleaning.py has no Streamlit import, so the logic can be tested and reused on its own. Run python test_cleaning.py to check it against the sample data.read_csv, .shape, .dtypes, .isnull(), .duplicated().fillna(), .dropna(), .drop_duplicates(), .astype(), .str.strip()np.percentile for IQR-based outlier boundsobject to a dedicated str type, so text columns are detected with pd.api.types.is_string_dtype(), which works on both versionscleaning.py already has a tested find_outliers_iqr() function that isn't wired into the UI yet. That's the next feature. After that:find_outliers_iqr, then drop or cap them"$1,200")Posted Oct 5, 2026
Cleaned CSV files, provided reports of changes made, and conducted data diagnosis.
0
0