Evaluation · Flagship experience

Probability Calibration & Thresholds

Is a predicted 0.8 probability actually an 80% event rate?

Start here

Is a predicted 0.8 probability actually an 80% event rate?

Discrimination asks whether positives rank above negatives; calibration asks whether predicted probabilities match observed frequencies. Thresholding then turns probabilities into decisions.

Building interactive view…
Understand

Build the mental model

Discrimination asks whether positives rank above negatives; calibration asks whether predicted probabilities match observed frequencies. Thresholding then turns probabilities into decisions. Separate model fitting, calibration and threshold choice. Choose thresholds from decision costs, not the default 0.5 by reflex.

What happens if…?

Break the assumption deliberately

Hold ranking fixed but distort probabilities, then compare AUC with calibration error.

Move the control and explain what you expect before reading the visual.

Technical lens

Formalise what the visual is doing

Reliability diagrams and Brier/log loss assess calibration. Platt, isotonic and temperature scaling transform scores using separate calibration evidence.

Technical questionUse a tiny case to make the mechanism observable. Reliability diagrams and Brier/log loss assess calibration. Platt, isotonic and temperature scaling transform scores using separate calibration evidence. Verify one intermediate quantity, state change or mapping independently; then predict the consequence of this change: Hold ranking fixed but distort probabilities, then compare AUC with calibration error.
Practitioner lens

Use it responsibly

Separate model fitting, calibration and threshold choice. Choose thresholds from decision costs, not the default 0.5 by reflex.

Transfer testUsing accuracy for a rare-event problem without checking class-specific errors.
Worked exploration

Use the visual as an experiment, not decoration

Collect cases predicted near 0.8. If only 55% are actually positive, the model is overconfident in that range. Compare a calibration curve with ROC-AUC: ranking and probability accuracy are different properties.

Technical lens

Reliability diagrams and Brier/log loss assess calibration. Platt, isotonic and temperature scaling transform scores using separate calibration evidence.

Practitioner check

Separate model fitting, calibration and threshold choice. Choose thresholds from decision costs, not the default 0.5 by reflex.

Prediction before interaction
Hold ranking fixed but distort probabilities, then compare AUC with calibration error.
Exploration walkthrough

Turn the interaction into an evidence trail

Collect cases predicted near 0.8. If only 55% are actually positive, the model is overconfident in that range. Compare a calibration curve with ROC-AUC: ranking and probability accuracy are different properties. Before moving the control, state your prediction. After the visual changes, name the specific state, statistic, boundary or mapping that changed and explain why that change is consistent—or inconsistent—with your prediction.

  • Record one observable quantity before the interaction and the same quantity afterwards.
  • Change one factor at a time so the causal effect of the control is inspectable.
  • Use an edge or failure case to discover where the concept stops behaving as the simple story suggests.
Visual demonstration of Probability Calibration & Thresholds
Static orientation diagram for Probability Calibration & Thresholds; use the interactive visual above to test how the relationships change.
Reference depth

Open the complete material

The flagship experience is the map. These pages contain the roads.

Continue this exact concept

Choose depth, practice or application.

These destinations are explicitly mapped to Probability Calibration & Thresholds; they are not generic landing-page fallbacks.