4 · Data Cleaning & Missing Data

Outliers & Anomalous Values

Outliers & Anomalous Values groups the core ideas a learner needs at the 4 · data cleaning & missing data stage. Work through the lessons in order when new to the area, or use them independently as a reference when implementing an analysis.

How to use this topic

Learn the mechanism one decision at a time

Work through the lessons in order if the topic is new. If you already know the basics, open the specific leaf lesson that matches the operation, diagnostic or failure mode you need.

1Definition→
2Mechanism→
3Example→
4Diagnostic→
5Decision
01
Data error vs legitimate extremeData error vs legitimate extreme is a practical concept within Outliers & Anomalous Values. It helps turn the broader workflow stage “4 · Data Cleaning & Missing Data” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.
02
IQR ruleIQR rule is a practical concept within Outliers & Anomalous Values. It helps turn the broader workflow stage “4 · Data Cleaning & Missing Data” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.
03
Z-score intuitionZ-score intuition is a practical concept within Outliers & Anomalous Values. It helps turn the broader workflow stage “4 · Data Cleaning & Missing Data” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.
04
Robust z-scores and MADRobust z-scores and MAD is a practical concept within Outliers & Anomalous Values. It helps turn the broader workflow stage “4 · Data Cleaning & Missing Data” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.
05
Winsorisation and clippingWinsorisation and clipping is a practical concept within Outliers & Anomalous Values. It helps turn the broader workflow stage “4 · Data Cleaning & Missing Data” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.
06
Transformations for skewTransformations for skew changes how raw variables are represented for analysis or modelling. The transformation should preserve the information needed by the task while making assumptions explicit and reproducible.
07
Model-based anomaly detectionModel-based anomaly detection represents a family or practice in model building. The central idea is to define what structure can be learned, how model quality is measured during fitting, and how generalisation is tested on observations not used to choose the model.
08
Documenting outlier decisionsDocumenting outlier decisions is a practical concept within Outliers & Anomalous Values. It helps turn the broader workflow stage “4 · Data Cleaning & Missing Data” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.