Capstone Project · Lesson 93

Evaluate

Final evaluation measures the frozen model on untouched test data using metrics chosen for the task.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

Evaluate

Final evaluation measures the frozen model on untouched test data using metrics chosen for the task. It should include diagnostic views when a scalar score could hide important errors.

Learning goal: explain why Evaluate behaves this way, apply it to a small example, and verify the result independently. Begin by being able to justify this first step: Predict once, calculate primary/secondary metrics, inspect confusion/residual/calibration/subgroup behaviour and save predictions.

Deeper walkthrough

Read Evaluate as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Predict once, calculate primary/secondary metrics, inspect confusion/residual/calibration/subgroup behaviour and save predictions. Stage 2: Keep the stage inside the same data/validation definitions used by the rest of the project. Stage 3: Save the evidence produced by this stage so the next stage can be audited.

Mechanism

Follow the transformation

Predict once, calculate primary/secondary metrics, inspect confusion/residual/calibration/subgroup behaviour and save predictions.

Keep the stage inside the same data/validation definitions used by the rest of the project.

Save the evidence produced by this stage so the next stage can be audited.

Evidence

Know what would convince you

  • Verify the split/validation boundary before comparing scores.
  • Inspect model/preprocessing state or a hand-computable tiny example.
Useful distinctionRepresentation: How the method encodes inputs/predictions.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Predict once

Predict once, calculate primary/secondary metrics, inspect confusion/residual/calibration/subgroup behaviour and save predictions.

Verification focus: record the evidence you inspected and the condition that would make this stage fail.
How it works

Trace the mechanism step by step

  1. Predict once, calculate primary/secondary metrics, inspect confusion/residual/calibration/subgroup behaviour and save predictions.
  2. Keep the stage inside the same data/validation definitions used by the rest of the project.
  3. Save the evidence produced by this stage so the next stage can be audited.
Worked demonstration

Evaluate evidence

Evaluate evidence
Evidence: all final metrics can be recomputed from the saved test prediction table.
Expected / illustrative result
The worked evidence makes the output of this project stage concrete and auditable.
Interpret the result.

For Evaluate, trace the specific input through the mechanism above and independently verify one returned value, state change or side effect.

Distinctions & related ideas

Place the concept correctly

RepresentationHow the method encodes inputs/predictions.
Learning/operationWhat fitted state or calculation changes.
ValidationIndependent evidence used to judge generalisation or correctness.
Use deliberately

When it is appropriate

Use Evaluate when it answers a defined question in Capstone Project and its inputs/assumptions match the current data or program state.

Boundary conditions

When to stop or reconsider

Reconsider Evaluate when the required information is unavailable, the operation would violate a validation/data boundary, or a simpler operation answers the question more transparently.

Common mistakes

Failure modes to recognise

  • Optimising on the final test set.
  • Ignoring feature scale/representation or split structure when the method depends on them.
  • Reporting a single score without checking errors, variance or operating conditions.
Verification

How to check the result

  • Verify the split/validation boundary before comparing scores.
  • Inspect model/preprocessing state or a hand-computable tiny example.
  • Perturb one input/hyperparameter and predict the expected direction or behaviour.
Hands-on practice

Demonstrate understanding

Try this:

Construct a tiny example of Evaluate. First predict once, calculate primary/secondary metrics, inspect confusion/residual/calibration/subgroup behaviour and save predictions. Then keep the stage inside the same data/validation definitions used by the rest of the project. Predict the result before execution and explain one boundary or failure case.

List the stage inputs and expected artifact, rerun it from a clean state, and compare against a concrete acceptance check.
Knowledge check

Check reasoning, not memorisation

Which approach best demonstrates understanding of Evaluate?

Quick reference

Remember the logic

Step 1Predict once, calculate primary/secondary metrics, inspect confusion/residual/calibration/subgroup behaviour and save predictions.
Step 2Keep the stage inside the same data/validation definitions used by the rest of the project.
Step 3Save the evidence produced by this stage so the next stage can be audited.
Lesson summary

What to remember

  • Final evaluation measures the frozen model on untouched test data using metrics chosen for the task. It should include diagnostic views when a scalar score could hide important errors.
  • Predict once, calculate primary/secondary metrics, inspect confusion/residual/calibration/subgroup behaviour and save predictions.
  • Optimising on the final test set.
  • Verify the split/validation boundary before comparing scores.