Predictive Modelling · Lesson 59

Logistic Regression

Logistic regression is a linear probabilistic classifier.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

Logistic Regression

Logistic regression is a linear probabilistic classifier. A weighted feature score is transformed by the sigmoid function into a probability between 0 and 1, then a separate decision threshold can convert that probability into a class label.

Learning goal: explain why Logistic Regression behaves this way, apply it to a small example, and verify the result independently. Begin by being able to justify this first step: Fit coefficients by minimising classification log loss, usually with regularisation.

Deeper walkthrough

Read Logistic Regression as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Fit coefficients by minimising classification log loss, usually with regularisation. Stage 2: Inspect predicted probabilities, not only hard classes. Stage 3: Choose thresholds using validation data and decision costs. Final checkpoint: Scale/encode features inside the pipeline when required.

Mechanism

Follow the transformation

Fit coefficients by minimising classification log loss, usually with regularisation.

Inspect predicted probabilities, not only hard classes.

Choose thresholds using validation data and decision costs.

Evidence

Know what would convince you

  • Confirm fitted transformations/models saw only training data.
  • Retain fold/test predictions so metrics can be recomputed independently.
Useful distinctionTraining evidence: Information allowed to influence fitted state.
Visual demonstration of Logistic Regression
Visual demonstration: use the diagram to trace the main objects and state changes involved in Logistic Regression.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Fit coefficients by minimising classification log…

Fit coefficients by minimising classification log loss, usually with regularisation. At this stage of Logistic Regression, keep the incoming data or object separate from the learned parameter, transformed object, or statistic so the change can be reproduced and independently checked.

Transformation focus: keep the input and produced parameters/result separate so the change is observable and reproducible.
How it works

Trace the mechanism step by step

  1. Fit coefficients by minimising classification log loss, usually with regularisation.
  2. Inspect predicted probabilities, not only hard classes.
  3. Choose thresholds using validation data and decision costs.
  4. Evaluate discrimination and calibration separately.
  5. Scale/encode features inside the pipeline when required.
Worked demonstration

Probability vs class

Probability = 0.72; threshold = 0.50 → positive.
Same probability; threshold = 0.80 → negative.
Expected / illustrative result
The fitted probability and the operational threshold are distinct parts of the system.
Interpret the result.

For Logistic Regression, connect the result to the fitted state, held-out data or prediction rule that produced it and independently check one prediction, split or metric component.

Distinctions & related ideas

Place the concept correctly

Training evidenceInformation allowed to influence fitted state.
Held-out evidenceIndependent observations used to estimate generalisation.
InterpretationWhat the result supports, with assumptions and limitations.
Use deliberately

When it is appropriate

Use Logistic Regression when it answers a defined question in Predictive Modelling and its inputs/assumptions match the current data or program state.

Boundary conditions

When to stop or reconsider

Reconsider Logistic Regression when the required information is unavailable, the operation would violate a validation/data boundary, or a simpler operation answers the question more transparently.

Common mistakes

Failure modes to recognise

  • Learning preprocessing/feature/model choices from held-out test information.
  • Comparing models under different splits or preprocessing and attributing the difference to the algorithm.
  • Turning an association or model explanation into an unsupported causal claim.
Verification

How to check the result

  • Confirm fitted transformations/models saw only training data.
  • Retain fold/test predictions so metrics can be recomputed independently.
  • Inspect errors/subgroups and compare with a baseline before generalising the conclusion.
Hands-on practice

Demonstrate understanding

Try this:

Construct a tiny example of Logistic Regression. First fit coefficients by minimising classification log loss, usually with regularisation. Then inspect predicted probabilities, not only hard classes. Predict the result before execution and explain one boundary or failure case.

Use a tiny fixed split or synthetic example. State what is fitted, what remains held out, and what result you expect before running it.
Knowledge check

Check reasoning, not memorisation

Which approach best demonstrates understanding of Logistic Regression?

Quick reference

Remember the logic

Step 1Fit coefficients by minimising classification log loss, usually with regularisation.
Step 2Inspect predicted probabilities, not only hard classes.
Step 3Choose thresholds using validation data and decision costs.
Step 4Evaluate discrimination and calibration separately.
Lesson summary

What to remember

  • Logistic regression is a linear probabilistic classifier. A weighted feature score is transformed by the sigmoid function into a probability between 0 and 1, then a separate decision threshold can convert that probability into a class label.
  • Fit coefficients by minimising classification log loss, usually with regularisation.
  • Learning preprocessing/feature/model choices from held-out test information.
  • Confirm fitted transformations/models saw only training data.