Linear Models · Lesson 25

Logistic Regression

Logistic regression is a linear probabilistic classifier.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Logistic Regression actually means

Logistic regression is a linear probabilistic classifier. It forms a linear score from the features and maps that score through the logistic function to a probability for the positive class. Multiclass softmax generalises the idea to a probability distribution across classes.

Logistic Regression matters because linear models provide an interpretable baseline and expose central ideas such as coefficient estimation, probability links and regularisation. They also show clearly how preprocessing and penalty choices affect fitted parameters.

Deeper walkthrough

Read Logistic Regression as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Compute linear score z = β0 + βᵀx. Stage 2: Map z through sigmoid p = 1/(1+e^-z) for binary classification. Stage 3: Fit coefficients by minimising log loss / maximising likelihood, often with regularisation. Final checkpoint: Evaluate discrimination and calibration separately.

Mechanism

Follow the transformation

Compute linear score z = β0 + βᵀx.

Map z through sigmoid p = 1/(1+e^-z) for binary classification.

Fit coefficients by minimising log loss / maximising likelihood, often with regularisation.

Evidence

Know what would convince you

  • Fit a tiny or baseline case first and confirm prediction shape/range and a few outputs.
  • Evaluate with the same held-out folds/metric as competing models and inspect variability, not just the mean.
Useful distinctionLinear regression: Continuous prediction with unbounded output.
Visual demonstration of Logistic Regression
Visual demonstration: use the diagram to trace the main objects and state changes involved in Logistic Regression.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Compute linear score z = β0…

Compute linear score z = β0 + βᵀx. At this stage of Logistic Regression, keep the incoming data or object separate from the learned parameter, transformed object, or statistic so the change can be reproduced and independently checked.

Transformation focus: keep the input and produced parameters/result separate so the change is observable and reproducible.
Mathematical / formal view
logit(p) = log(p/(1-p)) = β0 + βᵀx; therefore a one-unit feature change adds β_j to log-odds when other features are held fixed.
How it works

Trace the mechanism step by step

  1. Compute linear score z = β0 + βᵀx.
  2. Map z through sigmoid p = 1/(1+e^-z) for binary classification.
  3. Fit coefficients by minimising log loss / maximising likelihood, often with regularisation.
  4. Convert probability to a class using a threshold chosen for the decision context.
  5. Evaluate discrimination and calibration separately.
Worked demonstration

Make the concept concrete

Demonstration

Python / scikit-learn example

# Step 1 — Import the module so its functions/classes are available to the rest of this example.
import numpy as np
# Step 2 — Import only the named objects needed by the following steps, keeping dependencies explicit.
from sklearn.linear_model import LogisticRegression
# Step 3 — Construct `X` as an array so vectorised numerical operations can be applied consistently.
X=np.array([[0],[1],[2],[3],[4],[5]])
# Step 4 — Construct `y` as an array so vectorised numerical operations can be applied consistently.
y=np.array([0,0,0,1,1,1])
# Step 5 — Fit the model or transformer, learning its parameters from the supplied training data.
m=LogisticRegression().fit(X,y)
# Step 6 — Display the current value explicitly so the result/state can be inspected during execution.
print(np.round(m.predict_proba([[2.0],[4.0]])[:,1],3))
Expected / illustrative result
The probability at x=4 is higher than at x=2; a later threshold converts probabilities into class decisions.
Interpret the result.

For Logistic Regression, connect the reported result to the exact training/validation/prediction step that produced it and check one prediction, fold or metric component independently.

Distinctions & related ideas

Know what this is — and what it is not

Linear regressionContinuous prediction with unbounded output.
Logistic regressionBinary class probability via sigmoid.
Softmax regressionMulticlass probabilities summing to 1.
ThresholdDecision rule applied after probability estimation; not part of probability fitting itself.
Use deliberately

When it is appropriate

Use Logistic Regression when its inductive assumptions fit the feature/target structure and it can be compared fairly with a simpler baseline on unseen data.

Boundary conditions

When to stop or reconsider

Prefer a simpler or different model when the sample size, representation, computational budget, interpretability requirement or data geometry conflicts with this method.

Common mistakes

Failure modes to recognise

  • Judging the model only by training fit instead of generalisation on held-out data.
  • Comparing models with inconsistent preprocessing, folds or evaluation metrics.
  • Tuning complexity without checking a simple baseline, error patterns and variance across splits.
Verification

How to check the result

  • Fit a tiny or baseline case first and confirm prediction shape/range and a few outputs.
  • Evaluate with the same held-out folds/metric as competing models and inspect variability, not just the mean.
  • Inspect errors/residuals or decision boundaries and vary one key hyperparameter to verify expected behaviour.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Logistic Regression. First compute linear score z = β0 + βᵀx. Then map z through sigmoid p = 1/(1+e^-z) for binary classification. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Start with a small baseline and a fixed validation split/fold assignment. Predict what increasing or decreasing one complexity control should do before testing it.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Logistic Regression, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Compute linear score z = β0 + βᵀx.
Step 2Map z through sigmoid p = 1/(1+e^-z) for binary classification.
Step 3Fit coefficients by minimising log loss / maximising likelihood, often with regularisation.
Step 4Convert probability to a class using a threshold chosen for the decision context.
Lesson summary

What to remember

  • Logistic regression is a linear probabilistic classifier. It forms a linear score from the features and maps that score through the logistic function to a probability for the positive class. Multiclass softmax generalises the idea to a probability distribution across classes.
  • Compute linear score z = β0 + βᵀx.
  • Judging the model only by training fit instead of generalisation on held-out data.
  • Fit a tiny or baseline case first and confirm prediction shape/range and a few outputs.