4 · Data Cleaning & Missing Data

Data Cleaning & Quality Control

Correct or flag duplicate, inconsistent, impossible and noisy records while preserving an auditable trail from raw to cleaned data. The topic is split into focused lessons so definitions, implementation decisions, diagnostics and common failure modes can be learned separately.

How to use this topic

Learn the mechanism one decision at a time

Work through the lessons in order if the topic is new. If you already know the basics, open the specific leaf lesson that matches the operation, diagnostic or failure mode you need.

1Definition→
2Mechanism→
3Example→
4Diagnostic→
5Decision
01
Duplicates and entity resolutionDuplicate rows or repeated entities can distort frequencies and leak between train and test. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
02
Range and consistency rulesValidation rules identify values that violate known domains, units or cross-field constraints. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
03
Outlier handlingOutliers can be errors, rare valid cases or the most important events in the dataset. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
04
Label qualitySupervised models inherit ambiguity, noise and bias from target labels. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.
05
Cleaning reproducibilityCleaning rules should be executable, versioned and tested rather than applied manually. The important practical question is not only how the technique is defined, but what assumptions it introduces, which data are allowed to influence it, and how its effect should be validated on unseen evidence.