Evaluation & Diagnostics · Lesson 73

Subgroup Evaluation

Subgroup Evaluation is part of model evaluation.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Subgroup Evaluation actually means

Subgroup Evaluation is part of model evaluation. A metric is a compressed view of model behaviour, so reliable evaluation uses several complementary summaries plus plots and subgroup/error analysis.

Subgroup Evaluation matters because a single metric compresses model behaviour. Diagnostics, curves, residuals, calibration, thresholds and subgroup results reveal different failure modes that can lead to different deployment decisions.

Deeper walkthrough

Read Subgroup Evaluation as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Evaluate predictions on data not used to fit/tune the model. Stage 2: Choose metrics that match the target type and decision costs. Stage 3: Inspect distributions/residuals or threshold curves rather than one number. Final checkpoint: Attach uncertainty to performance estimates when sample size or variability matters.

Mechanism

Follow the transformation

Evaluate predictions on data not used to fit/tune the model.

Choose metrics that match the target type and decision costs.

Inspect distributions/residuals or threshold curves rather than one number.

Evidence

Know what would convince you

  • Recompute the metric from a tiny set of predictions/labels or residuals by hand.
  • Inspect the underlying confusion counts, curve points, residuals or probability bins that generate the summary.
Useful distinctionAccuracy: Share of correct labels; can hide minority-class failure.
How it works

Trace the mechanism step by step

  1. Evaluate predictions on data not used to fit/tune the model.
  2. Choose metrics that match the target type and decision costs.
  3. Inspect distributions/residuals or threshold curves rather than one number.
  4. Check subgroup performance and calibration when predictions drive decisions.
  5. Attach uncertainty to performance estimates when sample size or variability matters.
Worked demonstration

Make the concept concrete

Demonstration

Text example

Overall accuracy = 90%.
Group A accuracy = 94% (n=900).
Group B accuracy = 54% (n=100).
Expected / illustrative result
The aggregate hides a serious subgroup failure; report group size and uncertainty as well as performance.
Interpret the result.

For Subgroup Evaluation, connect the reported result to the exact training/validation/prediction step that produced it and check one prediction, fold or metric component independently.

Distinctions & related ideas

Know what this is — and what it is not

AccuracyShare of correct labels; can hide minority-class failure.
PrecisionAmong predicted positives, fraction truly positive.
RecallAmong actual positives, fraction detected.
ROC-AUCRanking across thresholds; may look optimistic under severe imbalance.
PR-AUCPrecision-recall trade-off; often more informative for rare positives.
MAE/RMSEAbsolute vs squared-error regression summaries.
Use deliberately

When it is appropriate

Use Subgroup Evaluation when the metric/diagnostic corresponds to the prediction type, positive class or business/scientific consequence that matters.

Boundary conditions

When to stop or reconsider

Do not compress performance to this measure alone when thresholds, class imbalance, calibration, subgroup behaviour or error costs change the decision.

Common mistakes

Failure modes to recognise

  • Using a metric without defining the positive class, averaging rule, units or threshold that gives it meaning.
  • Optimising/selecting on the final test set or changing the threshold after seeing test performance.
  • Reporting one scalar while hiding calibration, residual structure, class imbalance or subgroup failures.
Verification

How to check the result

  • Recompute the metric from a tiny set of predictions/labels or residuals by hand.
  • Inspect the underlying confusion counts, curve points, residuals or probability bins that generate the summary.
  • Evaluate the same frozen predictions across relevant thresholds/subgroups and explain which decision trade-off changes.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Subgroup Evaluation. First evaluate predictions on data not used to fit/tune the model. Then choose metrics that match the target type and decision costs. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Use 6–10 labelled predictions or residuals and calculate the measure manually. Keep the threshold/positive class/units explicit and connect the number to a decision.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Subgroup Evaluation, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Evaluate predictions on data not used to fit/tune the model.
Step 2Choose metrics that match the target type and decision costs.
Step 3Inspect distributions/residuals or threshold curves rather than one number.
Step 4Check subgroup performance and calibration when predictions drive decisions.
Lesson summary

What to remember

  • Subgroup Evaluation is part of model evaluation. A metric is a compressed view of model behaviour, so reliable evaluation uses several complementary summaries plus plots and subgroup/error analysis.
  • Evaluate predictions on data not used to fit/tune the model.
  • Using a metric without defining the positive class, averaging rule, units or threshold that gives it meaning.
  • Recompute the metric from a tiny set of predictions/labels or residuals by hand.