Evaluation & Diagnostics · Lesson 70

Calibration

Calibration is part of model evaluation.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Calibration actually means

Calibration is part of model evaluation. A metric is a compressed view of model behaviour, so reliable evaluation uses several complementary summaries plus plots and subgroup/error analysis.

Calibration matters because a single metric compresses model behaviour. Diagnostics, curves, residuals, calibration, thresholds and subgroup results reveal different failure modes that can lead to different deployment decisions.

Deeper walkthrough

Read Calibration as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Evaluate predictions on data not used to fit/tune the model. Stage 2: Choose metrics that match the target type and decision costs. Stage 3: Inspect distributions/residuals or threshold curves rather than one number. Final checkpoint: Attach uncertainty to performance estimates when sample size or variability matters.

Mechanism

Follow the transformation

Evaluate predictions on data not used to fit/tune the model.

Choose metrics that match the target type and decision costs.

Inspect distributions/residuals or threshold curves rather than one number.

Evidence

Know what would convince you

  • Recompute the metric from a tiny set of predictions/labels or residuals by hand.
  • Inspect the underlying confusion counts, curve points, residuals or probability bins that generate the summary.
Useful distinctionAccuracy: Share of correct labels; can hide minority-class failure.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Evaluate predictions on data not used…

Evaluate predictions on data not used to fit/tune the model. At this stage of Calibration, keep the incoming data or object separate from the learned parameter, transformed object, or statistic so the change can be reproduced and independently checked.

Transformation focus: keep the input and produced parameters/result separate so the change is observable and reproducible.
How it works

Trace the mechanism step by step

  1. Evaluate predictions on data not used to fit/tune the model.
  2. Choose metrics that match the target type and decision costs.
  3. Inspect distributions/residuals or threshold curves rather than one number.
  4. Check subgroup performance and calibration when predictions drive decisions.
  5. Attach uncertainty to performance estimates when sample size or variability matters.
Worked demonstration

Make the concept concrete

Demonstration

Text example

Among 100 cases assigned probability ≈0.70, a calibrated model should produce roughly 70 positives over repeated comparable cases.
Calibration asks whether probability magnitudes match observed frequencies; discrimination asks whether positives rank above negatives.
Expected / illustrative result
A model can rank well yet be systematically overconfident or underconfident.
Interpret the result.

For Calibration, connect the reported result to the exact training/validation/prediction step that produced it and check one prediction, fold or metric component independently.

Distinctions & related ideas

Know what this is — and what it is not

AccuracyShare of correct labels; can hide minority-class failure.
PrecisionAmong predicted positives, fraction truly positive.
RecallAmong actual positives, fraction detected.
ROC-AUCRanking across thresholds; may look optimistic under severe imbalance.
PR-AUCPrecision-recall trade-off; often more informative for rare positives.
MAE/RMSEAbsolute vs squared-error regression summaries.
Use deliberately

When it is appropriate

Use Calibration when the metric/diagnostic corresponds to the prediction type, positive class or business/scientific consequence that matters.

Boundary conditions

When to stop or reconsider

Do not compress performance to this measure alone when thresholds, class imbalance, calibration, subgroup behaviour or error costs change the decision.

Common mistakes

Failure modes to recognise

  • Using a metric without defining the positive class, averaging rule, units or threshold that gives it meaning.
  • Optimising/selecting on the final test set or changing the threshold after seeing test performance.
  • Reporting one scalar while hiding calibration, residual structure, class imbalance or subgroup failures.
Verification

How to check the result

  • Recompute the metric from a tiny set of predictions/labels or residuals by hand.
  • Inspect the underlying confusion counts, curve points, residuals or probability bins that generate the summary.
  • Evaluate the same frozen predictions across relevant thresholds/subgroups and explain which decision trade-off changes.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Calibration. First evaluate predictions on data not used to fit/tune the model. Then choose metrics that match the target type and decision costs. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Use 6–10 labelled predictions or residuals and calculate the measure manually. Keep the threshold/positive class/units explicit and connect the number to a decision.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Calibration, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Evaluate predictions on data not used to fit/tune the model.
Step 2Choose metrics that match the target type and decision costs.
Step 3Inspect distributions/residuals or threshold curves rather than one number.
Step 4Check subgroup performance and calibration when predictions drive decisions.
Lesson summary

What to remember

  • Calibration is part of model evaluation. A metric is a compressed view of model behaviour, so reliable evaluation uses several complementary summaries plus plots and subgroup/error analysis.
  • Evaluate predictions on data not used to fit/tune the model.
  • Using a metric without defining the positive class, averaging rule, units or threshold that gives it meaning.
  • Recompute the metric from a tiny set of predictions/labels or residuals by hand.