Is a predicted 0.8 probability actually an 80% event rate?
Start here
Is a predicted 0.8 probability actually an 80% event rate?
Discrimination asks whether positives rank above negatives; calibration asks whether predicted probabilities match observed frequencies. Thresholding then turns probabilities into decisions.
Building interactive view…
Understand
Build the mental model
Discrimination asks whether positives rank above negatives; calibration asks whether predicted probabilities match observed frequencies. Thresholding then turns probabilities into decisions. Separate model fitting, calibration and threshold choice. Choose thresholds from decision costs, not the default 0.5 by reflex.
What happens if…?
Break the assumption deliberately
Hold ranking fixed but distort probabilities, then compare AUC with calibration error.
Move the control and explain what you expect before reading the visual.
Technical lens
Formalise what the visual is doing
Reliability diagrams and Brier/log loss assess calibration. Platt, isotonic and temperature scaling transform scores using separate calibration evidence.
Technical questionUse a tiny case to make the mechanism observable. Reliability diagrams and Brier/log loss assess calibration. Platt, isotonic and temperature scaling transform scores using separate calibration evidence. Verify one intermediate quantity, state change or mapping independently; then predict the consequence of this change: Hold ranking fixed but distort probabilities, then compare AUC with calibration error.
Practitioner lens
Use it responsibly
Separate model fitting, calibration and threshold choice. Choose thresholds from decision costs, not the default 0.5 by reflex.
Transfer testUsing accuracy for a rare-event problem without checking class-specific errors.
Worked exploration
Use the visual as an experiment, not decoration
Collect cases predicted near 0.8. If only 55% are actually positive, the model is overconfident in that range. Compare a calibration curve with ROC-AUC: ranking and probability accuracy are different properties.
Technical lens
Reliability diagrams and Brier/log loss assess calibration. Platt, isotonic and temperature scaling transform scores using separate calibration evidence.
Practitioner check
Separate model fitting, calibration and threshold choice. Choose thresholds from decision costs, not the default 0.5 by reflex.
Prediction before interaction
Hold ranking fixed but distort probabilities, then compare AUC with calibration error.
Exploration walkthrough
Turn the interaction into an evidence trail
Collect cases predicted near 0.8. If only 55% are actually positive, the model is overconfident in that range. Compare a calibration curve with ROC-AUC: ranking and probability accuracy are different properties. Before moving the control, state your prediction. After the visual changes, name the specific state, statistic, boundary or mapping that changed and explain why that change is consistent—or inconsistent—with your prediction.
Record one observable quantity before the interaction and the same quantity afterwards.
Change one factor at a time so the causal effect of the control is inspectable.
Use an edge or failure case to discover where the concept stops behaving as the simple story suggests.
Static orientation diagram for Probability Calibration & Thresholds; use the interactive visual above to test how the relationships change.
Reference depth
Open the complete material
The flagship experience is the map. These pages contain the roads.