Start hereWhat changes after a model leaves the notebook?
Production models interact with changing data, systems and decisions. Monitoring asks whether inputs, predictions, calibration and outcomes still resemble the evidence used to validate the model.
Building interactive view…
Technical lensFormalise what the visual is doing
Data drift changes input distributions; concept drift changes P(y|x); performance drift is observed degradation. Detection requires reference windows, delayed labels and response policies.
Technical questionUse a tiny case to make the mechanism observable. Data drift changes input distributions; concept drift changes P(y|x); performance drift is observed degradation. Detection requires reference windows, delayed labels and response policies. Verify one intermediate quantity, state change or mapping independently; then predict the consequence of this change: Shift an input distribution without changing labels, then distinguish data drift from concept drift.
Practitioner lensUse it responsibly
Define thresholds and owners before deployment. Retraining should be governed, reproducible and evaluated against the current champion.
Transfer testTransfer this idea to a new example and justify each decision using this practitioner rule: Define thresholds and owners before deployment. Retraining should be governed, reproducible and evaluated against the current champion. Then explain what should change if you deliberately test: Shift an input distribution without changing labels, then distinguish data drift from concept drift.
Worked explorationUse the visual as an experiment, not decoration
Compare training feature distributions with current production data and delayed labels. A shift in input age or missingness is data drift; a change in the relationship between features and target is concept drift. Decide what monitoring should trigger investigation rather than automatic retraining.
Technical lens
Data drift changes input distributions; concept drift changes P(y|x); performance drift is observed degradation. Detection requires reference windows, delayed labels and response policies.
Practitioner check
Define thresholds and owners before deployment. Retraining should be governed, reproducible and evaluated against the current champion.
Prediction before interactionShift an input distribution without changing labels, then distinguish data drift from concept drift.
Exploration walkthroughTurn the interaction into an evidence trail
Compare training feature distributions with current production data and delayed labels. A shift in input age or missingness is data drift; a change in the relationship between features and target is concept drift. Decide what monitoring should trigger investigation rather than automatic retraining. Before moving the control, state your prediction. After the visual changes, name the specific state, statistic, boundary or mapping that changed and explain why that change is consistent—or inconsistent—with your prediction.
- Record one observable quantity before the interaction and the same quantity afterwards.
- Change one factor at a time so the causal effect of the control is inspectable.
- Use an edge or failure case to discover where the concept stops behaving as the simple story suggests.