Every analysis is only as reliable as the data behind it. This section documents data cleaning projects where raw, corrupted, and inconsistent datasets were systematically transformed into analysis-ready sources. Working across SQL and Python, each project tackled real-world data quality problems like placeholder values, wrong data types, missing fields, formatting inconsistencies, and unrecoverable records, using structured, reproducible pipelines. Clean data is not the end goal; trustworthy analysis is.
See how each dataset was cleaned and prepared in the projects below.