Data Quality & Cleaning
After this lesson, you will be able to:
- Spot the four biggest data problems: missing values, outliers (weird extreme values), duplicates, and inconsistencies
- Classify missingness as MCAR, MAR, or MNAR — because the type of missingness determines which imputation strategy is safe to use
- Pick the right fix for each problem — fill in blanks, remove bad rows, cap extreme values, or deduplicate — based on why the problem happened
- Build a step-by-step data cleaning process that keeps a record of every change so you can always explain or undo what you did, and measure quality with completeness, uniqueness, and validity scores
Before You Start
Don't worry if this feels tedious — cleaning data is not glamorous, but it is genuinely the highest-impact skill in all of ML. Once you get the hang of it, it becomes second nature!