Follow the transformation
Scale features when units differ.
Choose k and a distance metric using validation.
Compute distances from a query to training cases.
K-nearest neighbours predicts from labelled training examples closest to the query under a chosen distance metric.
K-nearest neighbours predicts from labelled training examples closest to the query under a chosen distance metric. It stores the training data rather than learning a compact global equation, so feature scale, k and local sample density directly affect predictions.
Learning goal: explain why KNN behaves this way, apply it to a small example, and verify the result independently. Begin by being able to justify this first step: Scale features when units differ.
Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Scale features when units differ. Stage 2: Choose k and a distance metric using validation. Stage 3: Compute distances from a query to training cases. Final checkpoint: Inspect sensitivity to k and high-dimensional/noisy features.
Scale features when units differ.
Choose k and a distance metric using validation.
Compute distances from a query to training cases.
Scale features when units differ. For KNN, identify the exact state before this stage, the operation or rule applied here, and the observable state afterwards so the mechanism remains inspectable.
Nearest labels to a query: [A, A, B], k=3 → prediction A.The result follows directly from the local neighbourhood; changing feature scaling can change which cases are nearest.
For KNN, connect the result to the fitted state, held-out data or prediction rule that produced it and independently check one prediction, split or metric component.
Training evidenceInformation allowed to influence fitted state.Held-out evidenceIndependent observations used to estimate generalisation.InterpretationWhat the result supports, with assumptions and limitations.Use KNN when it answers a defined question in Predictive Modelling and its inputs/assumptions match the current data or program state.
Reconsider KNN when the required information is unavailable, the operation would violate a validation/data boundary, or a simpler operation answers the question more transparently.
Construct a tiny example of KNN. First scale features when units differ. Then choose k and a distance metric using validation. Predict the result before execution and explain one boundary or failure case.
Which approach best demonstrates understanding of KNN?
Step 1Scale features when units differ.Step 2Choose k and a distance metric using validation.Step 3Compute distances from a query to training cases.Step 4Use neighbour voting for classification or averaging for regression.