Metrics, Calibration & Thresholds · Lesson 75

PR AUC

PR AUC evaluates a particular aspect of prediction quality.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What PR AUC actually means

PR AUC evaluates a particular aspect of prediction quality. No single metric captures discrimination, calibration, threshold costs and subgroup performance simultaneously; choose metrics based on the decision and class/base-rate structure.

PR AUC matters because classification and regression metrics answer different operational questions. Ranking, probability quality, threshold decisions and error magnitudes should be evaluated separately when the application cares about them separately.

Deeper walkthrough

Read PR AUC as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Generate predictions on held-out data. Stage 2: For probabilistic classifiers, separate probability quality from thresholded decisions. Stage 3: Compute metrics with the positive class and averaging convention explicitly defined. Final checkpoint: Choose an operating threshold using validation data and decision costs, then freeze it for test evaluation.

Mechanism

Follow the transformation

Generate predictions on held-out data.

For probabilistic classifiers, separate probability quality from thresholded decisions.

Compute metrics with the positive class and averaging convention explicitly defined.

Evidence

Know what would convince you

  • Recompute the metric from a tiny set of predictions/labels or residuals by hand.
  • Inspect the underlying confusion counts, curve points, residuals or probability bins that generate the summary.
Useful distinctionPrecision: TP/(TP+FP): reliability of positive predictions.
Visual demonstration of PR AUC
Visual demonstration: use the diagram to trace the main objects and state changes involved in PR AUC.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Generate predictions on held-out data

Generate predictions on held-out data. Treat the output from PR AUC as evidence to inspect: confirm its type, shape, range or units and connect it back to the input that produced it.

Output focus: inspect both the value and its shape/type/meaning before treating it as a trustworthy result.
How it works

Trace the mechanism step by step

  1. Generate predictions on held-out data.
  2. For probabilistic classifiers, separate probability quality from thresholded decisions.
  3. Compute metrics with the positive class and averaging convention explicitly defined.
  4. Inspect a curve/confusion matrix/residual distribution in addition to scalar metrics.
  5. Choose an operating threshold using validation data and decision costs, then freeze it for test evaluation.
Worked demonstration

Make the concept concrete

Demonstration

Text example

At stricter thresholds: precision often rises while recall falls.
At looser thresholds: recall rises while precision may fall.
PR curves focus on positive predictions and are strongly affected by prevalence.
Expected / illustrative result
Use the curve to select an operating region on validation data; evaluate the frozen threshold on held-out test data.
Interpret the result.

For PR AUC, connect the reported result to the exact training/validation/prediction step that produced it and check one prediction, fold or metric component independently.

Distinctions & related ideas

Know what this is — and what it is not

PrecisionTP/(TP+FP): reliability of positive predictions.
RecallTP/(TP+FN): detection of actual positives.
F1/FβHarmonic combination of precision/recall; β weights recall relative to precision.
ROC-AUCProbability a random positive ranks above a random negative.
PR-AUCPrecision-recall ranking summary, sensitive to prevalence.
Brier/log lossProbability-quality scores that penalise miscalibrated/confident errors.
Use deliberately

When it is appropriate

Use PR AUC when the metric/diagnostic corresponds to the prediction type, positive class or business/scientific consequence that matters.

Boundary conditions

When to stop or reconsider

Do not compress performance to this measure alone when thresholds, class imbalance, calibration, subgroup behaviour or error costs change the decision.

Common mistakes

Failure modes to recognise

  • Using a metric without defining the positive class, averaging rule, units or threshold that gives it meaning.
  • Optimising/selecting on the final test set or changing the threshold after seeing test performance.
  • Reporting one scalar while hiding calibration, residual structure, class imbalance or subgroup failures.
Verification

How to check the result

  • Recompute the metric from a tiny set of predictions/labels or residuals by hand.
  • Inspect the underlying confusion counts, curve points, residuals or probability bins that generate the summary.
  • Evaluate the same frozen predictions across relevant thresholds/subgroups and explain which decision trade-off changes.
Connection to feature importance

AUC-based screening under imbalance

For strongly imbalanced targets, ROC-AUC and PR-AUC answer different ranking questions. A feature-level AUC screen should be interpreted alongside class prevalence and validated in the final multivariate pipeline.

Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of PR AUC. First generate predictions on held-out data. Then for probabilistic classifiers, separate probability quality from thresholded decisions. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Use 6–10 labelled predictions or residuals and calculate the measure manually. Keep the threshold/positive class/units explicit and connect the number to a decision.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from PR AUC, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Generate predictions on held-out data.
Step 2For probabilistic classifiers, separate probability quality from thresholded decisions.
Step 3Compute metrics with the positive class and averaging convention explicitly defined.
Step 4Inspect a curve/confusion matrix/residual distribution in addition to scalar metrics.
Lesson summary

What to remember

  • PR AUC evaluates a particular aspect of prediction quality. No single metric captures discrimination, calibration, threshold costs and subgroup performance simultaneously; choose metrics based on the decision and class/base-rate structure.
  • Generate predictions on held-out data.
  • Using a metric without defining the positive class, averaging rule, units or threshold that gives it meaning.
  • Recompute the metric from a tiny set of predictions/labels or residuals by hand.