Cleaning & Imputation · Lesson 29

Complete Case Analysis

Complete Case Analysis concerns missing-data mechanisms and treatments.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Complete Case Analysis actually means

Complete Case Analysis concerns missing-data mechanisms and treatments. Missingness can be informative: MCAR means missingness is independent of observed/unobserved values, MAR allows dependence on observed values, and MNAR involves dependence on the unobserved value or other unobserved processes.

Complete Case Analysis matters because missingness changes both the available sample and the information contained in each feature. Imputation is a modelling decision whose assumptions and validation boundary can affect bias and predictive performance.

Deeper walkthrough

Read Complete Case Analysis as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Profile where and when values are missing. Stage 2: Compare observed characteristics between missing and non-missing groups. Stage 3: Choose a treatment whose assumptions are defensible. Final checkpoint: Evaluate whether imputation changes distributions, relationships and downstream model performance.

Mechanism

Follow the transformation

Profile where and when values are missing.

Compare observed characteristics between missing and non-missing groups.

Choose a treatment whose assumptions are defensible.

Evidence

Know what would convince you

  • Compare missingness, distributions, row counts and key constraints before and after the change.
  • Inspect representative changed rows and confirm the rule with domain/data documentation.
Useful distinctionComplete case: Uses only rows with required observed values; simple but can waste data/bias results.
Visual demonstration of Complete Case Analysis
Visual demonstration: use the diagram to trace the main objects and state changes involved in Complete Case Analysis.
How it works

Trace the mechanism step by step

  1. Profile where and when values are missing.
  2. Compare observed characteristics between missing and non-missing groups.
  3. Choose a treatment whose assumptions are defensible.
  4. Fit imputation parameters only on training data in predictive workflows.
  5. Evaluate whether imputation changes distributions, relationships and downstream model performance.
Worked demonstration

Make the concept concrete

Demonstration

Python / pandas example

# Step 1 — Import the module so its functions/classes are available to the rest of this example.
import pandas as pd
# Step 2 — Construct `df` as a tabular object with named columns for inspectable analysis.
df = pd.DataFrame({"x":[1,None,3],"y":[10,20,30]})
# Step 3 — Display the current value explicitly so the result/state can be inspected during execution.
print("original rows",len(df))
# Step 4 — Display the current value explicitly so the result/state can be inspected during execution.
print("complete rows",len(df.dropna()))
# Step 5 — Display the current value explicitly so the result/state can be inspected during execution.
print("median-imputed x",df.x.fillna(df.x.median()).tolist())
Expected / illustrative result
original rows 3; complete rows 2; median-imputed x [1.0, 2.0, 3.0]. The choice changes sample size and assumptions.
Interpret the result.

For Complete Case Analysis, trace representative source rows/columns into the result and reconcile row counts, dtypes, keys or missing values that the operation could change.

Distinctions & related ideas

Know what this is — and what it is not

Complete caseUses only rows with required observed values; simple but can waste data/bias results.
Simple imputationFixed statistic; stable but attenuates variability.
KNN / iterativeUses relationships with other features; stronger assumptions and more computation.
IndicatorAdds a feature marking missingness so models can learn associated structure.
Use deliberately

When it is appropriate

Use Complete Case Analysis when it helps diagnose, document or correct a data-quality issue without destroying information needed for the downstream question.

Boundary conditions

When to stop or reconsider

Do not “clean” automatically when the apparent anomaly may carry signal, reflect data collection, or require domain adjudication; preserve an audit trail of changes.

Common mistakes

Failure modes to recognise

  • Treating missing values, outliers or labels as purely technical defects without investigating how they were generated.
  • Applying learned cleaning/imputation using information from validation/test data.
  • Changing values without recording which rows changed and how distributions/counts were affected.
Verification

How to check the result

  • Compare missingness, distributions, row counts and key constraints before and after the change.
  • Inspect representative changed rows and confirm the rule with domain/data documentation.
  • In predictive work, fit learned cleaning only on training data and verify the pipeline reproduces that boundary.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Complete Case Analysis. First profile where and when values are missing. Then compare observed characteristics between missing and non-missing groups. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Start by measuring the problem, not fixing it. Keep a before/after table of counts or distributions and inspect the exact records affected by the rule.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Complete Case Analysis, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Profile where and when values are missing.
Step 2Compare observed characteristics between missing and non-missing groups.
Step 3Choose a treatment whose assumptions are defensible.
Step 4Fit imputation parameters only on training data in predictive workflows.
Lesson summary

What to remember

  • Complete Case Analysis concerns missing-data mechanisms and treatments. Missingness can be informative: MCAR means missingness is independent of observed/unobserved values, MAR allows dependence on observed values, and MNAR involves dependence on the unobserved value or other unobserved processes.
  • Profile where and when values are missing.
  • Treating missing values, outliers or labels as purely technical defects without investigating how they were generated.
  • Compare missingness, distributions, row counts and key constraints before and after the change.