Follow the transformation
Profile where and when values are missing.
Compare observed characteristics between missing and non-missing groups.
Choose a treatment whose assumptions are defensible.
Complete Case Analysis concerns missing-data mechanisms and treatments.
Complete Case Analysis concerns missing-data mechanisms and treatments. Missingness can be informative: MCAR means missingness is independent of observed/unobserved values, MAR allows dependence on observed values, and MNAR involves dependence on the unobserved value or other unobserved processes.
Complete Case Analysis matters because missingness changes both the available sample and the information contained in each feature. Imputation is a modelling decision whose assumptions and validation boundary can affect bias and predictive performance.
Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Profile where and when values are missing. Stage 2: Compare observed characteristics between missing and non-missing groups. Stage 3: Choose a treatment whose assumptions are defensible. Final checkpoint: Evaluate whether imputation changes distributions, relationships and downstream model performance.
Profile where and when values are missing.
Compare observed characteristics between missing and non-missing groups.
Choose a treatment whose assumptions are defensible.
# Step 1 — Import the module so its functions/classes are available to the rest of this example.
import pandas as pd
# Step 2 — Construct `df` as a tabular object with named columns for inspectable analysis.
df = pd.DataFrame({"x":[1,None,3],"y":[10,20,30]})
# Step 3 — Display the current value explicitly so the result/state can be inspected during execution.
print("original rows",len(df))
# Step 4 — Display the current value explicitly so the result/state can be inspected during execution.
print("complete rows",len(df.dropna()))
# Step 5 — Display the current value explicitly so the result/state can be inspected during execution.
print("median-imputed x",df.x.fillna(df.x.median()).tolist())original rows 3; complete rows 2; median-imputed x [1.0, 2.0, 3.0]. The choice changes sample size and assumptions.
For Complete Case Analysis, trace representative source rows/columns into the result and reconcile row counts, dtypes, keys or missing values that the operation could change.
Complete caseUses only rows with required observed values; simple but can waste data/bias results.Simple imputationFixed statistic; stable but attenuates variability.KNN / iterativeUses relationships with other features; stronger assumptions and more computation.IndicatorAdds a feature marking missingness so models can learn associated structure.Use Complete Case Analysis when it helps diagnose, document or correct a data-quality issue without destroying information needed for the downstream question.
Do not “clean” automatically when the apparent anomaly may carry signal, reflect data collection, or require domain adjudication; preserve an audit trail of changes.
Build a tiny, inspectable example of Complete Case Analysis. First profile where and when values are missing. Then compare observed characteristics between missing and non-missing groups. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.
Before trusting a result from Complete Case Analysis, which check provides the strongest evidence that you understand and applied it correctly?
Step 1Profile where and when values are missing.Step 2Compare observed characteristics between missing and non-missing groups.Step 3Choose a treatment whose assumptions are defensible.Step 4Fit imputation parameters only on training data in predictive workflows.