Metrics, Calibration & Thresholds · Lesson 71

Confusion Matrix

Confusion Matrix evaluates a particular aspect of prediction quality.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Confusion Matrix actually means

Confusion Matrix evaluates a particular aspect of prediction quality. No single metric captures discrimination, calibration, threshold costs and subgroup performance simultaneously; choose metrics based on the decision and class/base-rate structure.

Confusion Matrix matters because classification and regression metrics answer different operational questions. Ranking, probability quality, threshold decisions and error magnitudes should be evaluated separately when the application cares about them separately.

Deeper walkthrough

Read Confusion Matrix as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Generate predictions on held-out data. Stage 2: For probabilistic classifiers, separate probability quality from thresholded decisions. Stage 3: Compute metrics with the positive class and averaging convention explicitly defined. Final checkpoint: Choose an operating threshold using validation data and decision costs, then freeze it for test evaluation.

Mechanism

Follow the transformation

Generate predictions on held-out data.

For probabilistic classifiers, separate probability quality from thresholded decisions.

Compute metrics with the positive class and averaging convention explicitly defined.

Evidence

Know what would convince you

  • Recompute the metric from a tiny set of predictions/labels or residuals by hand.
  • Inspect the underlying confusion counts, curve points, residuals or probability bins that generate the summary.
Useful distinctionPrecision: TP/(TP+FP): reliability of positive predictions.
Visual demonstration of Confusion Matrix
Visual demonstration: use the diagram to trace the main objects and state changes involved in Confusion Matrix.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Generate predictions on held-out data

Generate predictions on held-out data. Treat the output from Confusion Matrix as evidence to inspect: confirm its type, shape, range or units and connect it back to the input that produced it.

Output focus: inspect both the value and its shape/type/meaning before treating it as a trustworthy result.
How it works

Trace the mechanism step by step

  1. Generate predictions on held-out data.
  2. For probabilistic classifiers, separate probability quality from thresholded decisions.
  3. Compute metrics with the positive class and averaging convention explicitly defined.
  4. Inspect a curve/confusion matrix/residual distribution in addition to scalar metrics.
  5. Choose an operating threshold using validation data and decision costs, then freeze it for test evaluation.
Worked demonstration

Make the concept concrete

Demonstration

Python example

# Step 1 — Compute the right-hand expression and store its result in `tp,fp,fn,tn` for the next step.
tp,fp,fn,tn = 30,10,5,55
# Step 2 — Compute the right-hand expression and store its result in `precision` for the next step.
precision=tp/(tp+fp)
# Step 3 — Compute the right-hand expression and store its result in `recall` for the next step.
recall=tp/(tp+fn)
# Step 4 — Compute the right-hand expression and store its result in `accuracy` for the next step.
accuracy=(tp+tn)/(tp+fp+fn+tn)
# Step 5 — Display the current value explicitly so the result/state can be inspected during execution.
print(round(accuracy,3), round(precision,3), round(recall,3))
Expected / illustrative result
accuracy=0.85, precision=0.75, recall≈0.857. The same confusion matrix supports multiple decision-relevant summaries.
Interpret the result.

For Confusion Matrix, connect the reported result to the exact training/validation/prediction step that produced it and check one prediction, fold or metric component independently.

Distinctions & related ideas

Know what this is — and what it is not

PrecisionTP/(TP+FP): reliability of positive predictions.
RecallTP/(TP+FN): detection of actual positives.
F1/FβHarmonic combination of precision/recall; β weights recall relative to precision.
ROC-AUCProbability a random positive ranks above a random negative.
PR-AUCPrecision-recall ranking summary, sensitive to prevalence.
Brier/log lossProbability-quality scores that penalise miscalibrated/confident errors.
Use deliberately

When it is appropriate

Use Confusion Matrix when the metric/diagnostic corresponds to the prediction type, positive class or business/scientific consequence that matters.

Boundary conditions

When to stop or reconsider

Do not compress performance to this measure alone when thresholds, class imbalance, calibration, subgroup behaviour or error costs change the decision.

Common mistakes

Failure modes to recognise

  • Using a metric without defining the positive class, averaging rule, units or threshold that gives it meaning.
  • Optimising/selecting on the final test set or changing the threshold after seeing test performance.
  • Reporting one scalar while hiding calibration, residual structure, class imbalance or subgroup failures.
Verification

How to check the result

  • Recompute the metric from a tiny set of predictions/labels or residuals by hand.
  • Inspect the underlying confusion counts, curve points, residuals or probability bins that generate the summary.
  • Evaluate the same frozen predictions across relevant thresholds/subgroups and explain which decision trade-off changes.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Confusion Matrix. First generate predictions on held-out data. Then for probabilistic classifiers, separate probability quality from thresholded decisions. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Use 6–10 labelled predictions or residuals and calculate the measure manually. Keep the threshold/positive class/units explicit and connect the number to a decision.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Confusion Matrix, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Generate predictions on held-out data.
Step 2For probabilistic classifiers, separate probability quality from thresholded decisions.
Step 3Compute metrics with the positive class and averaging convention explicitly defined.
Step 4Inspect a curve/confusion matrix/residual distribution in addition to scalar metrics.
Lesson summary

What to remember

  • Confusion Matrix evaluates a particular aspect of prediction quality. No single metric captures discrimination, calibration, threshold costs and subgroup performance simultaneously; choose metrics based on the decision and class/base-rate structure.
  • Generate predictions on held-out data.
  • Using a metric without defining the positive class, averaging rule, units or threshold that gives it meaning.
  • Recompute the metric from a tiny set of predictions/labels or residuals by hand.