Missing Data: Advanced Imputation
Missing Data: Advanced Imputation groups the core ideas a learner needs at the 4 · data cleaning & missing data stage. Work through the lessons in order when new to the area, or use them independently as a reference when implementing an analysis.
Learn the mechanism one decision at a time
Work through the lessons in order if the topic is new. If you already know the basics, open the specific leaf lesson that matches the operation, diagnostic or failure mode you need.
1Definition→
2Mechanism→
3Example→
4Diagnostic→
5Decision
KNN imputationKNN imputation replaces a missing feature using values from nearby observations, where “nearby” is computed from the other available features. It can preserve local nonlinear structure better than a global mean, but distance quality depends heavily on scale and irrelevant features.
02Iterative imputation / chained equationsIterative imputation treats each incomplete variable as a prediction problem. It cycles through features, predicts one feature from the others, updates the filled values and repeats. This captures multivariate relationships that simple univariate imputation ignores.
03Regression imputationRegression imputation is a practical concept within Missing Data: Advanced Imputation. It helps turn the broader workflow stage “4 · Data Cleaning & Missing Data” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.
04Multiple imputation intuitionMultiple imputation intuition is a practical concept within Missing Data: Advanced Imputation. It helps turn the broader workflow stage “4 · Data Cleaning & Missing Data” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.
05Missing-indicator featuresA missing indicator is a binary feature that records whether the original value was absent. Imputation supplies a usable numeric/categorical value; the indicator lets the model learn whether the fact of being missing itself carries information.
06Native missing-value handling in modelsNative missing-value handling in models represents a family or practice in model building. The central idea is to define what structure can be learned, how model quality is measured during fitting, and how generalisation is tested on observations not used to choose the model.
07Choosing an imputation strategyChoosing an imputation strategy is a practical concept within Missing Data: Advanced Imputation. It helps turn the broader workflow stage “4 · Data Cleaning & Missing Data” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.