Cleaning & Imputation · Lesson 28

MCAR MAR and MNAR

MCAR MAR and MNAR concerns missing-data mechanisms and treatments.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What MCAR MAR and MNAR actually means

MCAR MAR and MNAR concerns missing-data mechanisms and treatments. Missingness can be informative: MCAR means missingness is independent of observed/unobserved values, MAR allows dependence on observed values, and MNAR involves dependence on the unobserved value or other unobserved processes.

MCAR MAR and MNAR matters because missingness changes both the available sample and the information contained in each feature. Imputation is a modelling decision whose assumptions and validation boundary can affect bias and predictive performance.

Deeper walkthrough

Read MCAR MAR and MNAR as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Profile where and when values are missing. Stage 2: Compare observed characteristics between missing and non-missing groups. Stage 3: Choose a treatment whose assumptions are defensible. Final checkpoint: Evaluate whether imputation changes distributions, relationships and downstream model performance.

Mechanism

Follow the transformation

Profile where and when values are missing.

Compare observed characteristics between missing and non-missing groups.

Choose a treatment whose assumptions are defensible.

Evidence

Know what would convince you

  • Compare missingness, distributions, row counts and key constraints before and after the change.
  • Inspect representative changed rows and confirm the rule with domain/data documentation.
Useful distinctionComplete case: Uses only rows with required observed values; simple but can waste data/bias results.
How it works

Trace the mechanism step by step

  1. Profile where and when values are missing.
  2. Compare observed characteristics between missing and non-missing groups.
  3. Choose a treatment whose assumptions are defensible.
  4. Fit imputation parameters only on training data in predictive workflows.
  5. Evaluate whether imputation changes distributions, relationships and downstream model performance.
Worked demonstration

Make the concept concrete

Demonstration

Text example

MCAR example: a random scanner failure removes values independently of patient/data properties.
MAR example: income is more often missing among younger respondents, and age is observed.
MNAR example: very high earners are less likely to report income specifically because the unreported income is high.
Expected / illustrative result
The mechanism is a statement about why values are missing, not a property that can always be proven from the observed table alone.
Interpret the result.

For MCAR MAR and MNAR, trace representative source rows/columns into the result and reconcile row counts, dtypes, keys or missing values that the operation could change.

Distinctions & related ideas

Know what this is — and what it is not

Complete caseUses only rows with required observed values; simple but can waste data/bias results.
Simple imputationFixed statistic; stable but attenuates variability.
KNN / iterativeUses relationships with other features; stronger assumptions and more computation.
IndicatorAdds a feature marking missingness so models can learn associated structure.
Use deliberately

When it is appropriate

Use MCAR MAR and MNAR when it helps diagnose, document or correct a data-quality issue without destroying information needed for the downstream question.

Boundary conditions

When to stop or reconsider

Do not “clean” automatically when the apparent anomaly may carry signal, reflect data collection, or require domain adjudication; preserve an audit trail of changes.

Common mistakes

Failure modes to recognise

  • Treating missing values, outliers or labels as purely technical defects without investigating how they were generated.
  • Applying learned cleaning/imputation using information from validation/test data.
  • Changing values without recording which rows changed and how distributions/counts were affected.
Verification

How to check the result

  • Compare missingness, distributions, row counts and key constraints before and after the change.
  • Inspect representative changed rows and confirm the rule with domain/data documentation.
  • In predictive work, fit learned cleaning only on training data and verify the pipeline reproduces that boundary.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of MCAR MAR and MNAR. First profile where and when values are missing. Then compare observed characteristics between missing and non-missing groups. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Start by measuring the problem, not fixing it. Keep a before/after table of counts or distributions and inspect the exact records affected by the rule.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from MCAR MAR and MNAR, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Profile where and when values are missing.
Step 2Compare observed characteristics between missing and non-missing groups.
Step 3Choose a treatment whose assumptions are defensible.
Step 4Fit imputation parameters only on training data in predictive workflows.
Lesson summary

What to remember

  • MCAR MAR and MNAR concerns missing-data mechanisms and treatments. Missingness can be informative: MCAR means missingness is independent of observed/unobserved values, MAR allows dependence on observed values, and MNAR involves dependence on the unobserved value or other unobserved processes.
  • Profile where and when values are missing.
  • Treating missing values, outliers or labels as purely technical defects without investigating how they were generated.
  • Compare missingness, distributions, row counts and key constraints before and after the change.