Reference the entire data-to-decision workflow.
Use this as an encyclopedia: start with foundations, then follow the workflow through collection, ingestion, EDA, cleaning, feature engineering, validation, modelling, evaluation, interpretation, reporting and monitoring. Every leaf lesson includes explanation, visuals, practical guidance and code.
Browse by workflow stage or search a concept.
01Foundations10 topics · 68 lessonsWorkflow stage
Analytics, Data Science & AI LandscapeUnderstand how data analytics, business intelligence, data science, artificial intelligence, machine learning, deep learning, data engineering and MLOps overlap and differ. In this8
Open topic overview →
Data AnalyticsData analytics turns raw observations into evidence for questions, decisions and operational improvement. It spans descriptive, diagnostic, Business IntelligenceBusiness intelligence focuses on governed reporting, dashboards, semantic metrics and repeatable decision support, often over enterprise datData ScienceData science combines statistics, programming, domain knowledge and experimental reasoning to extract insight and build data-driven systems.Artificial IntelligenceArtificial intelligence is the broad field of building systems that perform tasks associated with perception, reasoning, planning, language,Machine LearningMachine learning learns reusable patterns from examples or interaction rather than encoding every decision rule manually. The important pracDeep LearningDeep learning uses multi-layer neural networks to learn hierarchical representations, often directly from high-dimensional raw data. The impData EngineeringData engineering builds reliable systems for ingesting, transforming, storing and serving data so analytics and ML receive reproducible inpuMLOpsMLOps applies software, data and operational engineering practices to the lifecycle of machine-learning systems. The important practical que
Common Data Science & ML TasksConnect business questions to common analytical task types such as classification, regression, clustering, anomaly detection, dimensionality reduction and forecasting. In this topi6
Open topic overview →
ClassificationClassification predicts one of a finite set of classes, often with probabilities or scores before a final class decision. The important pracRegressionRegression predicts a continuous numeric quantity. The important practical question is not only how the technique is defined, but what assumClusteringClustering groups observations according to similarity or density without known class labels. The important practical question is not only hAnomaly detectionAnomaly detection identifies observations that are unusual relative to normal behaviour or the learned data distribution. The important pracDimensionality reductionDimensionality reduction maps many input variables into a smaller representation while preserving selected structure. The important practicaForecastingForecasting predicts future values while respecting temporal order, trend, seasonality and changing relationships. The important practical q
Learning ParadigmsLearn the major ways models receive supervision or feedback: supervised, unsupervised, semi-supervised, self-supervised, reinforcement and generative learning. In this topic is exp6
Open topic overview →
Supervised learningSupervised learning trains on input-output pairs where a target label or numeric outcome is known. The important practical question is not oUnsupervised learningUnsupervised learning discovers structure without an explicit target variable. The important practical question is not only how the techniquSemi-supervised learningSemi-supervised learning combines a small labelled set with a larger unlabelled set. The important practical question is not only how the teSelf-supervised learningSelf-supervised learning creates training signals from the data itself, such as masked-token prediction, contrastive pairs or reconstructionReinforcement learningReinforcement learning learns actions through interaction and delayed rewards. The important practical question is not only how the techniquGenerative learningGenerative learning models the distribution or structure of data so new samples, sequences or representations can be produced. The important
Mathematical & Statistical FoundationsReview the mathematical ideas that appear repeatedly in analytics and machine learning: probability, statistics, linear algebra, optimisation, loss functions and generalisation. In5
Open topic overview →
Statistics and probabilityStatistics summarises data and quantifies uncertainty; probability provides a language for uncertain events and distributions. The importantLinear algebraLinear algebra describes vectors, matrices, projections and transformations used throughout ML. The important practical question is not onlyOptimisationOptimisation searches for parameters or decisions that minimise loss or maximise utility subject to constraints. The important practical queLoss and objective functionsA loss measures how undesirable a prediction is; an objective may combine loss with regularisation or other constraints. The important practGeneralisation, bias and varianceGeneralisation is performance on unseen data. Bias reflects systematic underfitting; variance reflects sensitivity to training data. The imp
Types of AnalyticsTypes of Analytics groups the core ideas a learner needs at the foundations stage. Work through the lessons in order when new to the area, or use them independently as a reference 5
Open topic overview →
Descriptive analyticsDescriptive analytics is a practical concept within Types of Analytics. It helps turn the broader workflow stage “Foundations” into an expliDiagnostic analyticsDiagnostic analytics is a practical concept within Types of Analytics. It helps turn the broader workflow stage “Foundations” into an explicPredictive analyticsPredictive analytics is a practical concept within Types of Analytics. It helps turn the broader workflow stage “Foundations” into an explicPrescriptive analyticsPrescriptive analytics is a practical concept within Types of Analytics. It helps turn the broader workflow stage “Foundations” into an explExploratory vs confirmatory analysisExploratory vs confirmatory analysis is a practical concept within Types of Analytics. It helps turn the broader workflow stage “Foundations
Data Types, Measurement & VariablesData Types, Measurement & Variables groups the core ideas a learner needs at the foundations stage. Work through the lessons in order when new to the area, or use them independentl6
Open topic overview →
Numeric, categorical and boolean variablesNumeric, categorical and boolean variables is a practical concept within Data Types, Measurement & Variables. It helps turn the broader workNominal, ordinal, interval and ratio scalesNominal, ordinal, interval and ratio scales changes how raw variables are represented for analysis or modelling. The transformation should pDiscrete vs continuous variablesDiscrete vs continuous variables is a practical concept within Data Types, Measurement & Variables. It helps turn the broader workflow stageIdentifiers, keys and metadataIdentifiers, keys and metadata is a core data-integration operation. In analytics, correctness depends not only on syntax but on the relatioFeatures, targets and labelsFeatures, targets and labels changes how raw variables are represented for analysis or modelling. The transformation should preserve the infWide vs long dataWide vs long data is a practical concept within Data Types, Measurement & Variables. It helps turn the broader workflow stage “Foundations”
Python, Notebooks & Reproducible AnalysisPython, Notebooks & Reproducible Analysis groups the core ideas a learner needs at the foundations stage. Work through the lessons in order when new to the area, or use them indepe6
Open topic overview →
Python objects and variablesPython objects and variables is a practical concept within Python, Notebooks & Reproducible Analysis. It helps turn the broader workflow staLists, dictionaries and arraysLists, dictionaries and arrays is a practical concept within Python, Notebooks & Reproducible Analysis. It helps turn the broader workflow sFunctions and reusable codeFunctions and reusable code is a practical concept within Python, Notebooks & Reproducible Analysis. It helps turn the broader workflow stagJupyter notebook workflowJupyter notebook workflow is a practical concept within Python, Notebooks & Reproducible Analysis. It helps turn the broader workflow stage Virtual environments and dependenciesVirtual environments and dependencies is a practical concept within Python, Notebooks & Reproducible Analysis. It helps turn the broader worRandom seeds and reproducibilityRandom seeds and reproducibility is a practical concept within Python, Notebooks & Reproducible Analysis. It helps turn the broader workflow
SQL & Relational Data BasicsSQL & Relational Data Basics groups the core ideas a learner needs at the foundations stage. Work through the lessons in order when new to the area, or use them independently as a 6
Open topic overview →
Tables, rows, columns and keysTables, rows, columns and keys is a core data-integration operation. In analytics, correctness depends not only on syntax but on the relatioSELECT, WHERE and ORDER BYSELECT, WHERE and ORDER BY is a practical concept within SQL & Relational Data Basics. It helps turn the broader workflow stage “FoundationsGROUP BY and aggregationGROUP BY and aggregation is a core data-integration operation. In analytics, correctness depends not only on syntax but on the relationship JOIN fundamentalsJOIN fundamentals is a core data-integration operation. In analytics, correctness depends not only on syntax but on the relationship betweenWindow functionsWindow functions is a core data-integration operation. In analytics, correctness depends not only on syntax but on the relationship between CTEs and readable analytical SQLCTEs and readable analytical SQL is a core data-integration operation. In analytics, correctness depends not only on syntax but on the relat
Visualisation FoundationsVisualisation Foundations groups the core ideas a learner needs at the foundations stage. Work through the lessons in order when new to the area, or use them independently as a ref6
Open topic overview →
Choose chart by analytical questionChoose chart by analytical question is a communication and diagnostic technique that maps data or results into a visual form. A useful visuaPosition, length, area and colour encodingsPosition, length, area and colour encodings changes how raw variables are represented for analysis or modelling. The transformation should pBar, line, scatter and distribution plotsBar, line, scatter and distribution plots is a communication and diagnostic technique that maps data or results into a visual form. A usefulSmall multiples and facetingSmall multiples and faceting is a practical concept within Visualisation Foundations. It helps turn the broader workflow stage “Foundations”Labels, scales and annotationsLabels, scales and annotations changes how raw variables are represented for analysis or modelling. The transformation should preserve the iAvoid misleading visualisationsAvoid misleading visualisations is a communication and diagnostic technique that maps data or results into a visual form. A useful visual ma
Excel & Spreadsheet AnalyticsUse spreadsheets as auditable analytical tools: structure data as tables, write transparent formulas, perform lookups and conditional aggregation, build PivotTables, create readabl14
Open topic overview →
Workbook, worksheet and table anatomyUnderstand workbooks, worksheets, ranges, rows, columns and Excel Tables before writing formulas.Cell references: relative, absolute and mixedLearn how A1, $A$1, A$1 and $A1 behave when formulas are copied.Core arithmetic and aggregation formulasUse SUM, AVERAGE, MIN, MAX and COUNT to build transparent summaries.Logical formulas with IF, IFS, AND and ORTranslate business rules into readable spreadsheet logic.Conditional aggregation with SUMIFS, COUNTIFS and AVERAGEIFSCalculate metrics for selected segments without manual filtering.Lookup analysis with XLOOKUPRetrieve attributes from a reference table using an exact key match.INDEX and MATCH for flexible lookupsUnderstand a composable lookup pattern and why key uniqueness matters.Text cleaning with TRIM, CLEAN, TEXTSPLIT and SUBSTITUTEStandardise messy text before grouping or joining.Date and time analysis in ExcelBuild month, quarter and elapsed-time fields from real Excel dates.Excel Tables and structured referencesUse table names and structured formulas so analysis expands safely with new rows.PivotTables for grouped analysisSummarise a table by dimensions and measures, then verify totals against source data.Charts in Excel: choose, label and auditCreate readable charts from a clean analytical table and avoid misleading axes.Power Query fundamentalsImport, clean and reshape data with a reproducible query instead of repeated manual edits.Spreadsheet error handling and auditingUse IFERROR carefully, trace precedents and build reconciliation checks.
021 · Problem Framing & Data Collection6 topics · 34 lessonsWorkflow stage
Data Collection, Sampling & SourcesAcquire data that represent the target population and deployment process, with known provenance, sampling logic and measurement quality. In this topic is expanded into smaller less5
Open topic overview →
Primary and secondary dataPrimary data are collected specifically for the project; secondary data are reused from existing systems, studies or public sources. The impSampling strategiesSampling determines which units from a population enter the dataset. The important practical question is not only how the technique is definMeasurement and instrumentationReliable measurements require calibrated instruments, stable definitions and timestamps. The important practical question is not only how thAPIs, files, databases and streamsData can arrive through static files, database queries, APIs, event streams or object storage. The important practical question is not only Dataset versioning and lineageLineage records how a dataset was produced from source inputs and transformations. The important practical question is not only how the tech
Data Governance, Privacy & EthicsUse data lawfully, proportionately and transparently while protecting privacy, security and affected populations. In this topic is expanded into smaller lessons so that definitions5
Open topic overview →
Data minimisationCollect and retain only information necessary for the stated purpose. The important practical question is not only how the technique is defiPrivacy and de-identificationPrivacy controls reduce the risk that individuals can be identified or their information misused. The important practical question is not onConsent, purpose and accessData use should align with consent, lawful authority, organisational policy and purpose limitations. The important practical question is notFairness and representativenessEvaluate whether data coverage and system behaviour differ systematically across relevant groups. The important practical question is not onDocumentation and accountabilityDocument datasets, assumptions, intended use, limitations, approvals and model decisions. The important practical question is not only how t
Problem Framing & Analytical DesignTranslate an organisational or scientific question into a precise analytical problem with a target, unit of analysis, decision context and success criterion. In this topic is expan5
Open topic overview →
Define the decision or research questionStart with the decision, hypothesis or action the analysis must support, not with an available dataset. The important practical question is Define the unit of analysisThe unit of analysis is the entity represented by one prediction or observation at decision time. The important practical question is not onConstruct the targetTarget construction converts the real-world outcome into a measurable label or value with a clear observation window. The important practicaChoose success criteriaTechnical metrics should be connected to practical cost, risk, capacity or scientific objectives. The important practical question is not onPlan the experimentDecide how data will be split, what baselines are needed, which comparisons are fair and what information must remain blinded. The important
Business & Research Problem FramingBusiness & Research Problem Framing groups the core ideas a learner needs at the 1 · problem framing & data collection stage. Work through the lessons in order when new to the area6
Open topic overview →
Stakeholders and decision contextStakeholders and decision context is a practical concept within Business & Research Problem Framing. It helps turn the broader workflow stagTranslate a question into an analytical taskTranslate a question into an analytical task is a practical concept within Business & Research Problem Framing. It helps turn the broader woDefine population and scopeDefine population and scope is a practical concept within Business & Research Problem Framing. It helps turn the broader workflow stage “1 ·Define KPIs and success metricsDefine KPIs and success metrics is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usePredictive vs causal questionsPredictive vs causal questions is a practical concept within Business & Research Problem Framing. It helps turn the broader workflow stage “Assumptions and decision costsAssumptions and decision costs is a practical concept within Business & Research Problem Framing. It helps turn the broader workflow stage “
Data Acquisition MethodsData Acquisition Methods groups the core ideas a learner needs at the 1 · problem framing & data collection stage. Work through the lessons in order when new to the area, or use th6
Open topic overview →
Surveys and questionnairesSurveys and questionnaires is a practical concept within Data Acquisition Methods. It helps turn the broader workflow stage “1 · Problem FraExperiments and A/B testsExperiments and A/B tests is a practical concept within Data Acquisition Methods. It helps turn the broader workflow stage “1 · Problem FramObservational dataObservational data is a practical concept within Data Acquisition Methods. It helps turn the broader workflow stage “1 · Problem Framing & DSensors and IoT dataSensors and IoT data is a practical concept within Data Acquisition Methods. It helps turn the broader workflow stage “1 · Problem Framing &Web and API data collectionWeb and API data collection belongs to the operational phase where an analytical result becomes a maintained system. Production quality requLogs and event instrumentationLogs and event instrumentation is a practical concept within Data Acquisition Methods. It helps turn the broader workflow stage “1 · Problem
Sampling & Study DesignSampling & Study Design groups the core ideas a learner needs at the 1 · problem framing & data collection stage. Work through the lessons in order when new to the area, or use the7
Open topic overview →
Simple random samplingSimple random sampling is part of study design: it determines which units enter the dataset and therefore which population the analysis can Stratified samplingStratified sampling is part of study design: it determines which units enter the dataset and therefore which population the analysis can legCluster and multistage samplingCluster and multistage sampling is part of study design: it determines which units enter the dataset and therefore which population the analSystematic samplingSystematic sampling is part of study design: it determines which units enter the dataset and therefore which population the analysis can legConvenience and purposive samplingConvenience and purposive sampling is part of study design: it determines which units enter the dataset and therefore which population the aSampling bias and coverage errorSampling bias and coverage error is part of study design: it determines which units enter the dataset and therefore which population the anaSample size and power intuitionSample size and power intuition is part of study design: it determines which units enter the dataset and therefore which population the anal
032 · Data Ingestion, Storage & Integration3 topics · 19 lessonsWorkflow stage
Files, Formats & Data IngestionFiles, Formats & Data Ingestion groups the core ideas a learner needs at the 2 · data ingestion, storage & integration stage. Work through the lessons in order when new to the area6
Open topic overview →
CSV and delimited textCSV and delimited text is a practical concept within Files, Formats & Data Ingestion. It helps turn the broader workflow stage “2 · Data IngExcel workbooksExcel workbooks is a practical concept within Files, Formats & Data Ingestion. It helps turn the broader workflow stage “2 · Data Ingestion,JSON and nested dataJSON and nested data is a practical concept within Files, Formats & Data Ingestion. It helps turn the broader workflow stage “2 · Data IngesParquet and columnar formatsParquet and columnar formats is a practical concept within Files, Formats & Data Ingestion. It helps turn the broader workflow stage “2 · DaReading data safely with explicit dtypesReading data safely with explicit dtypes is a practical concept within Files, Formats & Data Ingestion. It helps turn the broader workflow sChunked and large-file ingestionChunked and large-file ingestion is a practical concept within Files, Formats & Data Ingestion. It helps turn the broader workflow stage “2
Relational Data & JoinsRelational Data & Joins groups the core ideas a learner needs at the 2 · data ingestion, storage & integration stage. Work through the lessons in order when new to the area, or use7
Open topic overview →
Primary and foreign keysPrimary and foreign keys is a core data-integration operation. In analytics, correctness depends not only on syntax but on the relationship One-to-one joinsOne-to-one joins is a core data-integration operation. In analytics, correctness depends not only on syntax but on the relationship between One-to-many joinsOne-to-many joins is a core data-integration operation. In analytics, correctness depends not only on syntax but on the relationship betweenMany-to-many joinsMany-to-many joins is a core data-integration operation. In analytics, correctness depends not only on syntax but on the relationship betweeInner, left, right and full joinsInner, left, right and full joins is a core data-integration operation. In analytics, correctness depends not only on syntax but on the relaJoin validation and row-count checksJoin validation and row-count checks is a core data-integration operation. In analytics, correctness depends not only on syntax but on the rReshaping with pivot and meltReshaping with pivot and melt belongs to the operational phase where an analytical result becomes a maintained system. Production quality re
Warehouses, Lakes & PipelinesWarehouses, Lakes & Pipelines groups the core ideas a learner needs at the 2 · data ingestion, storage & integration stage. Work through the lessons in order when new to the area, 6
Open topic overview →
Operational databases vs analytical warehousesOperational databases vs analytical warehouses is a practical concept within Warehouses, Lakes & Pipelines. It helps turn the broader workflStar schemas: facts and dimensionsStar schemas: facts and dimensions is a practical concept within Warehouses, Lakes & Pipelines. It helps turn the broader workflow stage “2 Data lakes and lakehousesData lakes and lakehouses is a practical concept within Warehouses, Lakes & Pipelines. It helps turn the broader workflow stage “2 · Data InETL vs ELTETL vs ELT is a practical concept within Warehouses, Lakes & Pipelines. It helps turn the broader workflow stage “2 · Data Ingestion, StoragBatch vs streaming pipelinesBatch vs streaming pipelines changes how raw variables are represented for analysis or modelling. The transformation should preserve the infData contracts and schema evolutionData contracts and schema evolution is a practical concept within Warehouses, Lakes & Pipelines. It helps turn the broader workflow stage “2
043 · Data Understanding & EDA4 topics · 24 lessonsWorkflow stage
Data Understanding & Exploratory AnalysisInspect structure, distributions, relationships, missingness and anomalies before choosing transformations or models. In this topic is expanded into smaller lessons so that definit5
Open topic overview →
Schema and data typesA schema describes variables, types, units, allowed ranges, keys and relationships. The important practical question is not only how the tecDescriptive statisticsDescriptive summaries quantify centre, spread, frequency and shape. The important practical question is not only how the technique is defineDistribution visualisationHistograms, density plots, boxplots and empirical CDFs reveal shape, tails and outliers. The important practical question is not only how thRelationships and correlationPairwise plots, correlations and grouped summaries explore associations among variables. The important practical question is not only how thMissingness profilingMissing data patterns can contain information about collection processes and bias. The important practical question is not only how the tech
Univariate ExplorationUnivariate Exploration groups the core ideas a learner needs at the 3 · data understanding & eda stage. Work through the lessons in order when new to the area, or use them independ6
Open topic overview →
Frequency tablesFrequency tables is a communication and diagnostic technique that maps data or results into a visual form. A useful visual makes comparisonsMeasures of centreMeasures of centre is a practical concept within Univariate Exploration. It helps turn the broader workflow stage “3 · Data Understanding & Measures of spreadMeasures of spread is a practical concept within Univariate Exploration. It helps turn the broader workflow stage “3 · Data Understanding & Quantiles and percentilesQuantiles and percentiles is a practical concept within Univariate Exploration. It helps turn the broader workflow stage “3 · Data UnderstanHistograms and density plotsHistograms and density plots is a communication and diagnostic technique that maps data or results into a visual form. A useful visual makesBoxplots and robust summariesBoxplots and robust summaries is a communication and diagnostic technique that maps data or results into a visual form. A useful visual make
Bivariate & Multivariate ExplorationBivariate & Multivariate Exploration groups the core ideas a learner needs at the 3 · data understanding & eda stage. Work through the lessons in order when new to the area, or use6
Open topic overview →
Scatterplots and associationScatterplots and association is a communication and diagnostic technique that maps data or results into a visual form. A useful visual makesCross-tabulation and contingency tablesCross-tabulation and contingency tables is a communication and diagnostic technique that maps data or results into a visual form. A useful vGrouped summariesGrouped summaries is a practical concept within Bivariate & Multivariate Exploration. It helps turn the broader workflow stage “3 · Data UndCorrelation matricesCorrelation matrices is a practical concept within Bivariate & Multivariate Exploration. It helps turn the broader workflow stage “3 · Data Pair plots and multivariate patternsPair plots and multivariate patterns is a communication and diagnostic technique that maps data or results into a visual form. A useful visuConfounding and Simpson’s paradoxConfounding and Simpson’s paradox is a practical concept within Bivariate & Multivariate Exploration. It helps turn the broader workflow sta
Data Quality ProfilingData Quality Profiling groups the core ideas a learner needs at the 3 · data understanding & eda stage. Work through the lessons in order when new to the area, or use them independ7
Open topic overview →
CompletenessCompleteness is a practical concept within Data Quality Profiling. It helps turn the broader workflow stage “3 · Data Understanding & EDA” iUniquenessUniqueness is a practical concept within Data Quality Profiling. It helps turn the broader workflow stage “3 · Data Understanding & EDA” intValidityValidity is a practical concept within Data Quality Profiling. It helps turn the broader workflow stage “3 · Data Understanding & EDA” into ConsistencyConsistency is a practical concept within Data Quality Profiling. It helps turn the broader workflow stage “3 · Data Understanding & EDA” inAccuracy and plausibilityAccuracy and plausibility is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulnesTimelinessTimeliness is a practical concept within Data Quality Profiling. It helps turn the broader workflow stage “3 · Data Understanding & EDA” intAutomated profiling checksAutomated profiling checks is a practical concept within Data Quality Profiling. It helps turn the broader workflow stage “3 · Data Understa
054 · Data Cleaning & Missing Data10 topics · 65 lessonsWorkflow stage
Data Cleaning & Quality ControlCorrect or flag duplicate, inconsistent, impossible and noisy records while preserving an auditable trail from raw to cleaned data. In this topic is expanded into smaller lessons s5
Open topic overview →
Duplicates and entity resolutionDuplicate rows or repeated entities can distort frequencies and leak between train and test. The important practical question is not only hoRange and consistency rulesValidation rules identify values that violate known domains, units or cross-field constraints. The important practical question is not only Outlier handlingOutliers can be errors, rare valid cases or the most important events in the dataset. The important practical question is not only how the tLabel qualitySupervised models inherit ambiguity, noise and bias from target labels. The important practical question is not only how the technique is deCleaning reproducibilityCleaning rules should be executable, versioned and tested rather than applied manually. The important practical question is not only how the
Data PreprocessingPreprocessing converts raw variables into a representation that models can learn from without accidentally exposing validation/test information. In this topic is expanded into smal5
Open topic overview →
Missing valuesImpute using statistics or learned models fitted on training data only. Missingness indicators can be useful when absence itself is informatScalingStandardisation centers/scales variance; Min–Max maps ranges; Robust scaling relies on quantiles and is less sensitive to extremes. The impoCategorical encodingOne-hot encoding is transparent for nominal variables. Ordinal encoding requires real order. Target encoding must be cross-fitted to avoid lTransformationsLog, Box–Cox/Yeo–Johnson and rank transformations can stabilise skew or variance but change interpretation. The important practical questionPipelinesPackage preprocessing and modelling into one pipeline so each validation fold learns transformations only from its training subset. The impo
Imbalanced LearningImbalanced learning addresses rare classes through evaluation, sampling, weighting, thresholding and suitable objectives rather than simply maximising overall accuracy. In this top5
Open topic overview →
Start with metricsUse class-wise recall, precision, PR-AUC, balanced accuracy and confusion matrices before changing the data. The important practical questioClass weightingWeighted losses increase the contribution of minority examples without inventing synthetic observations. The important practical question isResamplingRandom under/over-sampling and methods such as SMOTE change the training distribution. Resampling must happen inside each training fold. TheThresholdingA good ranking model may need a different decision threshold to satisfy sensitivity, precision, workload or cost constraints. The important Rare-event validationEnsure every fold contains enough minority cases and preserve groups/time where required; otherwise estimates become unstable or optimistic.
Missing Data: Concepts & DiagnosisMissing Data: Concepts & Diagnosis groups the core ideas a learner needs at the 4 · data cleaning & missing data stage. Work through the lessons in order when new to the area, or u7
Open topic overview →
Missing-value representationsMissing values are not a single universal token. In pandas and NumPy they can appear as `NaN`, `pd.NA`, `NaT` or domain-specific sentinel coMissingness rate by row and columnMissingness rate by row and column is a practical concept within Missing Data: Concepts & Diagnosis. It helps turn the broader workflow stagMissingness patterns and matricesMissingness patterns and matrices is a practical concept within Missing Data: Concepts & Diagnosis. It helps turn the broader workflow stageMCAR: missing completely at randomMCAR means the probability that a value is missing does not depend on observed or unobserved values. Under true MCAR, complete-case analysisMAR: missing at randomMAR means missingness may depend on variables you observed, but after conditioning on those observed variables it does not additionally depeMNAR: missing not at randomMNAR means the probability of missingness still depends on the unobserved value even after conditioning on observed data. Standard imputatioWhy missingness mechanism mattersWhy missingness mechanism matters is a practical concept within Missing Data: Concepts & Diagnosis. It helps turn the broader workflow stage
Missing Data: Deletion & Simple ImputationMissing Data: Deletion & Simple Imputation groups the core ideas a learner needs at the 4 · data cleaning & missing data stage. Work through the lessons in order when new to the ar7
Open topic overview →
Complete-case deletionComplete-case deletion is a practical concept within Missing Data: Deletion & Simple Imputation. It helps turn the broader workflow stage “4Pairwise deletionPairwise deletion is a practical concept within Missing Data: Deletion & Simple Imputation. It helps turn the broader workflow stage “4 · DaConstant-value imputationConstant-value imputation is a practical concept within Missing Data: Deletion & Simple Imputation. It helps turn the broader workflow stageMean imputationMean imputation is a practical concept within Missing Data: Deletion & Simple Imputation. It helps turn the broader workflow stage “4 · DataMedian imputationMedian imputation is a practical concept within Missing Data: Deletion & Simple Imputation. It helps turn the broader workflow stage “4 · DaMode / most-frequent imputationMode / most-frequent imputation is a practical concept within Missing Data: Deletion & Simple Imputation. It helps turn the broader workflowGroup-wise imputationGroup-wise imputation is a practical concept within Missing Data: Deletion & Simple Imputation. It helps turn the broader workflow stage “4
Missing Data: Advanced ImputationMissing Data: Advanced Imputation groups the core ideas a learner needs at the 4 · data cleaning & missing data stage. Work through the lessons in order when new to the area, or us7
Open topic overview →
KNN imputationKNN imputation replaces a missing feature using values from nearby observations, where “nearby” is computed from the other available featureIterative imputation / chained equationsIterative imputation treats each incomplete variable as a prediction problem. It cycles through features, predicts one feature from the otheRegression imputationRegression imputation is a practical concept within Missing Data: Advanced Imputation. It helps turn the broader workflow stage “4 · Data ClMultiple imputation intuitionMultiple imputation intuition is a practical concept within Missing Data: Advanced Imputation. It helps turn the broader workflow stage “4 ·Missing-indicator featuresA missing indicator is a binary feature that records whether the original value was absent. Imputation supplies a usable numeric/categoricalNative missing-value handling in modelsNative missing-value handling in models represents a family or practice in model building. The central idea is to define what structure can Choosing an imputation strategyChoosing an imputation strategy is a practical concept within Missing Data: Advanced Imputation. It helps turn the broader workflow stage “4
Missing Data: Time Series & DiagnosticsMissing Data: Time Series & Diagnostics groups the core ideas a learner needs at the 4 · data cleaning & missing data stage. Work through the lessons in order when new to the area,7
Open topic overview →
Forward fillForward fill is a practical concept within Missing Data: Time Series & Diagnostics. It helps turn the broader workflow stage “4 · Data CleanBackward fillBackward fill is a practical concept within Missing Data: Time Series & Diagnostics. It helps turn the broader workflow stage “4 · Data CleaLinear interpolationLinear interpolation estimates a missing value between two observed points by assuming a straight-line change across the gap. It is common iTime-aware interpolationTime-aware interpolation is a practical concept within Missing Data: Time Series & Diagnostics. It helps turn the broader workflow stage “4 Interpolation limits and gapsInterpolation limits and gaps is a practical concept within Missing Data: Time Series & Diagnostics. It helps turn the broader workflow stagValidate imputation distributionsValidate imputation distributions is a practical concept within Missing Data: Time Series & Diagnostics. It helps turn the broader workflow Compare model performance across imputation strategiesCompare model performance across imputation strategies represents a family or practice in model building. The central idea is to define what
Duplicates, Entities & ConsistencyDuplicates, Entities & Consistency groups the core ideas a learner needs at the 4 · data cleaning & missing data stage. Work through the lessons in order when new to the area, or u7
Open topic overview →
Exact duplicate rowsExact duplicate rows is a practical concept within Duplicates, Entities & Consistency. It helps turn the broader workflow stage “4 · Data ClDuplicate identifiersDuplicate identifiers is a practical concept within Duplicates, Entities & Consistency. It helps turn the broader workflow stage “4 · Data CFuzzy duplicate recordsFuzzy duplicate records is a practical concept within Duplicates, Entities & Consistency. It helps turn the broader workflow stage “4 · DataEntity resolutionEntity resolution is a practical concept within Duplicates, Entities & Consistency. It helps turn the broader workflow stage “4 · Data CleanCanonical categories and spellingCanonical categories and spelling is a practical concept within Duplicates, Entities & Consistency. It helps turn the broader workflow stageUnit and currency consistencyUnit and currency consistency is a practical concept within Duplicates, Entities & Consistency. It helps turn the broader workflow stage “4 Cross-field consistency rulesCross-field consistency rules is a practical concept within Duplicates, Entities & Consistency. It helps turn the broader workflow stage “4
Outliers & Anomalous ValuesOutliers & Anomalous Values groups the core ideas a learner needs at the 4 · data cleaning & missing data stage. Work through the lessons in order when new to the area, or use them8
Open topic overview →
Data error vs legitimate extremeData error vs legitimate extreme is a practical concept within Outliers & Anomalous Values. It helps turn the broader workflow stage “4 · DaIQR ruleIQR rule is a practical concept within Outliers & Anomalous Values. It helps turn the broader workflow stage “4 · Data Cleaning & Missing DaZ-score intuitionZ-score intuition is a practical concept within Outliers & Anomalous Values. It helps turn the broader workflow stage “4 · Data Cleaning & MRobust z-scores and MADRobust z-scores and MAD is a practical concept within Outliers & Anomalous Values. It helps turn the broader workflow stage “4 · Data CleaniWinsorisation and clippingWinsorisation and clipping is a practical concept within Outliers & Anomalous Values. It helps turn the broader workflow stage “4 · Data CleTransformations for skewTransformations for skew changes how raw variables are represented for analysis or modelling. The transformation should preserve the informaModel-based anomaly detectionModel-based anomaly detection represents a family or practice in model building. The central idea is to define what structure can be learnedDocumenting outlier decisionsDocumenting outlier decisions is a practical concept within Outliers & Anomalous Values. It helps turn the broader workflow stage “4 · Data
Text, Date & Category CleaningText, Date & Category Cleaning groups the core ideas a learner needs at the 4 · data cleaning & missing data stage. Work through the lessons in order when new to the area, or use t7
Open topic overview →
Whitespace and casingWhitespace and casing is a practical concept within Text, Date & Category Cleaning. It helps turn the broader workflow stage “4 · Data CleanString parsing and regular expressionsString parsing and regular expressions is a practical concept within Text, Date & Category Cleaning. It helps turn the broader workflow stagDate parsing and time zonesDate parsing and time zones is a practical concept within Text, Date & Category Cleaning. It helps turn the broader workflow stage “4 · DataCategory normalizationCategory normalization is a practical concept within Text, Date & Category Cleaning. It helps turn the broader workflow stage “4 · Data CleaRare categoriesRare categories is a practical concept within Text, Date & Category Cleaning. It helps turn the broader workflow stage “4 · Data Cleaning & Invalid codes and sentinelsInvalid codes and sentinels is a practical concept within Text, Date & Category Cleaning. It helps turn the broader workflow stage “4 · DataUnicode and encoding issuesUnicode and encoding issues changes how raw variables are represented for analysis or modelling. The transformation should preserve the info
065 · Data Preprocessing & Feature Engineering10 topics · 68 lessonsWorkflow stage
Feature Engineering & SelectionFeature engineering creates informative representations; feature selection removes redundant, noisy or costly variables. Both should be evaluated inside the validation process. In 5
Open topic overview →
Domain featuresRatios, interactions, lags, rolling statistics and physically meaningful transformations can expose structure that simple learners otherwiseFilter methodsVariance filters, correlation, mutual information and univariate tests rank features without repeatedly fitting the final model. The importaWrapper methodsRecursive Feature Elimination and sequential selection evaluate subsets using a predictive model, often at high computational cost. The impoEmbedded methodsL1 regularisation, tree importance and sparsity-inducing learners perform selection during model fitting. The important practical question iStability and leakageA selected feature should remain useful across folds and time. Selection performed before CV can leak target information. The important prac
Numeric PreprocessingNumeric Preprocessing groups the core ideas a learner needs at the 5 · data preprocessing & feature engineering stage. Work through the lessons in order when new to the area, or us8
Open topic overview →
StandardisationStandardisation is a practical concept within Numeric Preprocessing. It helps turn the broader workflow stage “5 · Data Preprocessing & FeatMin-max scalingMin-max scaling is a practical concept within Numeric Preprocessing. It helps turn the broader workflow stage “5 · Data Preprocessing & FeatRobust scalingRobust scaling is a practical concept within Numeric Preprocessing. It helps turn the broader workflow stage “5 · Data Preprocessing & FeatuMaxAbs scalingMaxAbs scaling is a practical concept within Numeric Preprocessing. It helps turn the broader workflow stage “5 · Data Preprocessing & FeatuLog transformsLog transforms changes how raw variables are represented for analysis or modelling. The transformation should preserve the information needePower transformsPower transforms changes how raw variables are represented for analysis or modelling. The transformation should preserve the information neeQuantile transformsQuantile transforms changes how raw variables are represented for analysis or modelling. The transformation should preserve the information When scaling is unnecessaryWhen scaling is unnecessary is a practical concept within Numeric Preprocessing. It helps turn the broader workflow stage “5 · Data Preproce
Categorical PreprocessingCategorical Preprocessing groups the core ideas a learner needs at the 5 · data preprocessing & feature engineering stage. Work through the lessons in order when new to the area, o7
Open topic overview →
One-hot encodingOne-hot encoding changes how raw variables are represented for analysis or modelling. The transformation should preserve the information neeOrdinal encodingOrdinal encoding changes how raw variables are represented for analysis or modelling. The transformation should preserve the information neeTarget encodingTarget encoding changes how raw variables are represented for analysis or modelling. The transformation should preserve the information needFrequency / count encodingFrequency / count encoding changes how raw variables are represented for analysis or modelling. The transformation should preserve the inforHigh-cardinality categoriesHigh-cardinality categories is a practical concept within Categorical Preprocessing. It helps turn the broader workflow stage “5 · Data PrepUnknown categories at inferenceUnknown categories at inference is a practical concept within Categorical Preprocessing. It helps turn the broader workflow stage “5 · Data Avoiding target leakage in encodersAvoiding target leakage in encoders is a practical concept within Categorical Preprocessing. It helps turn the broader workflow stage “5 · D
Datetime & Time-Series FeaturesDatetime & Time-Series Features groups the core ideas a learner needs at the 5 · data preprocessing & feature engineering stage. Work through the lessons in order when new to the a7
Open topic overview →
Calendar featuresCalendar features changes how raw variables are represented for analysis or modelling. The transformation should preserve the information neElapsed-time featuresElapsed-time features changes how raw variables are represented for analysis or modelling. The transformation should preserve the informatioLag featuresLag features changes how raw variables are represented for analysis or modelling. The transformation should preserve the information needed Rolling-window featuresRolling-window features changes how raw variables are represented for analysis or modelling. The transformation should preserve the informatExpanding-window featuresExpanding-window features changes how raw variables are represented for analysis or modelling. The transformation should preserve the informSeasonal encodingsSeasonal encodings changes how raw variables are represented for analysis or modelling. The transformation should preserve the information nLeakage-safe time featuresLeakage-safe time features changes how raw variables are represented for analysis or modelling. The transformation should preserve the infor
Text Feature PreparationText Feature Preparation groups the core ideas a learner needs at the 5 · data preprocessing & feature engineering stage. Work through the lessons in order when new to the area, or7
Open topic overview →
Tokenisation intuitionTokenisation intuition is a practical concept within Text Feature Preparation. It helps turn the broader workflow stage “5 · Data PreprocessBag-of-wordsBag-of-words is a practical concept within Text Feature Preparation. It helps turn the broader workflow stage “5 · Data Preprocessing & FeatTF-IDFTF-IDF is a practical concept within Text Feature Preparation. It helps turn the broader workflow stage “5 · Data Preprocessing & Feature EnN-gramsN-grams is a practical concept within Text Feature Preparation. It helps turn the broader workflow stage “5 · Data Preprocessing & Feature EText normalizationText normalization is a practical concept within Text Feature Preparation. It helps turn the broader workflow stage “5 · Data Preprocessing EmbeddingsEmbeddings is a practical concept within Text Feature Preparation. It helps turn the broader workflow stage “5 · Data Preprocessing & FeaturTrain-test vocabulary leakageTrain-test vocabulary leakage is a practical concept within Text Feature Preparation. It helps turn the broader workflow stage “5 · Data Pre
Image & Signal PreparationImage & Signal Preparation groups the core ideas a learner needs at the 5 · data preprocessing & feature engineering stage. Work through the lessons in order when new to the area, 6
Open topic overview →
Resize and resampleResize and resample is a practical concept within Image & Signal Preparation. It helps turn the broader workflow stage “5 · Data PreprocessiNormalisationNormalisation is a practical concept within Image & Signal Preparation. It helps turn the broader workflow stage “5 · Data Preprocessing & FData augmentationData augmentation is a practical concept within Image & Signal Preparation. It helps turn the broader workflow stage “5 · Data PreprocessingWindowing and segmentationWindowing and segmentation is a practical concept within Image & Signal Preparation. It helps turn the broader workflow stage “5 · Data PrepFrequency-domain featuresFrequency-domain features changes how raw variables are represented for analysis or modelling. The transformation should preserve the informTrain-only augmentation rulesTrain-only augmentation rules is a practical concept within Image & Signal Preparation. It helps turn the broader workflow stage “5 · Data P
Feature ConstructionFeature Construction groups the core ideas a learner needs at the 5 · data preprocessing & feature engineering stage. Work through the lessons in order when new to the area, or use7
Open topic overview →
Ratios and ratesRatios and rates is a practical concept within Feature Construction. It helps turn the broader workflow stage “5 · Data Preprocessing & FeatInteractionsInteractions is a practical concept within Feature Construction. It helps turn the broader workflow stage “5 · Data Preprocessing & Feature Polynomial featuresPolynomial features changes how raw variables are represented for analysis or modelling. The transformation should preserve the information Binning and discretisationBinning and discretisation is a practical concept within Feature Construction. It helps turn the broader workflow stage “5 · Data PreprocessDomain aggregatesDomain aggregates is a practical concept within Feature Construction. It helps turn the broader workflow stage “5 · Data Preprocessing & FeaGroup-level featuresGroup-level features changes how raw variables are represented for analysis or modelling. The transformation should preserve the informationFeature stores and reuseFeature stores and reuse changes how raw variables are represented for analysis or modelling. The transformation should preserve the informa
Feature SelectionFeature Selection groups the core ideas a learner needs at the 5 · data preprocessing & feature engineering stage. Work through the lessons in order when new to the area, or use th7
Open topic overview →
Variance filteringVariance filtering is a practical concept within Feature Selection. It helps turn the broader workflow stage “5 · Data Preprocessing & FeatuUnivariate statistical testsUnivariate statistical tests is a practical concept within Feature Selection. It helps turn the broader workflow stage “5 · Data PreprocessiMutual informationMutual information is a practical concept within Feature Selection. It helps turn the broader workflow stage “5 · Data Preprocessing & FeatuRecursive feature eliminationRecursive feature elimination changes how raw variables are represented for analysis or modelling. The transformation should preserve the inSequential feature selectionSequential feature selection changes how raw variables are represented for analysis or modelling. The transformation should preserve the infL1 and model-based selectionL1 and model-based selection represents a family or practice in model building. The central idea is to define what structure can be learned,RFECV and nested selectionRFECV and nested selection is a practical concept within Feature Selection. It helps turn the broader workflow stage “5 · Data Preprocessing
Dimensionality ReductionDimensionality Reduction groups the core ideas a learner needs at the 5 · data preprocessing & feature engineering stage. Work through the lessons in order when new to the area, or7
Open topic overview →
PCA intuitionPCA intuition is a practical concept within Dimensionality Reduction. It helps turn the broader workflow stage “5 · Data Preprocessing & FeaChoosing number of componentsChoosing number of components is a practical concept within Dimensionality Reduction. It helps turn the broader workflow stage “5 · Data PreExplained varianceExplained variance is a practical concept within Dimensionality Reduction. It helps turn the broader workflow stage “5 · Data Preprocessing Truncated SVDTruncated SVD is a practical concept within Dimensionality Reduction. It helps turn the broader workflow stage “5 · Data Preprocessing & Feat-SNE for visualisationt-SNE for visualisation is a communication and diagnostic technique that maps data or results into a visual form. A useful visual makes compUMAP for visualisationUMAP for visualisation is a communication and diagnostic technique that maps data or results into a visual form. A useful visual makes compaFit transforms inside training foldsFit transforms inside training folds changes how raw variables are represented for analysis or modelling. The transformation should preserve
Composite PipelinesComposite Pipelines groups the core ideas a learner needs at the 5 · data preprocessing & feature engineering stage. Work through the lessons in order when new to the area, or use 7
Open topic overview →
Pipeline conceptPipeline concept changes how raw variables are represented for analysis or modelling. The transformation should preserve the information neeColumnTransformer`ColumnTransformer` applies different preprocessing pipelines to different subsets of columns and concatenates the resulting feature matrix.Different transforms by column typeDifferent transforms by column type changes how raw variables are represented for analysis or modelling. The transformation should preserve Custom transformersCustom transformers changes how raw variables are represented for analysis or modelling. The transformation should preserve the information Transformed targetsTransformed targets changes how raw variables are represented for analysis or modelling. The transformation should preserve the information Pipeline parameter searchPipeline parameter search changes how raw variables are represented for analysis or modelling. The transformation should preserve the informPersisting the fitted pipelinePersisting the fitted pipeline changes how raw variables are represented for analysis or modelling. The transformation should preserve the i
076 · Splitting, Validation & Experiment Design6 topics · 40 lessonsWorkflow stage
Cross-Validation & Data SplittingValidation estimates how a model will behave on unseen data. The split strategy must mirror the real independence structure of samples, subjects, groups and time. In this topic is 5
Open topic overview →
Holdout and K-FoldHoldout reserves one test partition. K-Fold rotates the validation partition through K disjoint folds so every observation is validated onceStratificationStratified splitting preserves class proportions. It is valuable for classification, especially when minority classes are small, but it doesGrouped validationGroup K-Fold keeps all records from the same subject, patient, customer, device or site together. Stratified Group K-Fold additionally triesTime-aware validationTime-Series Split, rolling windows and expanding windows preserve chronological order. Future observations must never influence training feaNested cross-validationNested CV separates hyperparameter selection from performance estimation: an inner loop tunes the model and an outer loop estimates generali
Data Leakage & Experimental IntegrityData leakage occurs when information unavailable at real prediction time influences training, feature construction, selection, tuning or evaluation. In this topic is expanded into 5
Open topic overview →
Preprocessing leakageScaling, imputation, PCA or feature selection fitted on the full dataset exposes validation/test distribution information. The important praTarget leakageFeatures directly or indirectly encode outcomes that would not be known at prediction time. The important practical question is not only howGroup leakageRecords from the same patient, customer, device, recording or source appear in both train and test sets. The important practical question isTemporal leakageFuture observations influence historical features, normalization, imputation or model tuning. The important practical question is not only hDuplicate and pipeline leakageNear-duplicates, augmented copies or cached derived features can cross split boundaries unless lineage is tracked. The important practical q
Train / Validation / Test DesignTrain / Validation / Test Design groups the core ideas a learner needs at the 6 · splitting, validation & experiment design stage. Work through the lessons in order when new to the7
Open topic overview →
Why separate training and evaluationWhy separate training and evaluation is a practical concept within Train / Validation / Test Design. It helps turn the broader workflow stagHoldout splitHoldout split is a practical concept within Train / Validation / Test Design. It helps turn the broader workflow stage “6 · Splitting, ValidThree-way train-validation-testThree-way train-validation-test is a practical concept within Train / Validation / Test Design. It helps turn the broader workflow stage “6 Stratified splitStratified split is a practical concept within Train / Validation / Test Design. It helps turn the broader workflow stage “6 · Splitting, VaGroup-aware splitGroup-aware split is a practical concept within Train / Validation / Test Design. It helps turn the broader workflow stage “6 · Splitting, VTime-aware splitTime-aware split is a practical concept within Train / Validation / Test Design. It helps turn the broader workflow stage “6 · Splitting, VaSeed stabilitySeed stability is a practical concept within Train / Validation / Test Design. It helps turn the broader workflow stage “6 · Splitting, Vali
Cross-Validation MethodsCross-Validation Methods groups the core ideas a learner needs at the 6 · splitting, validation & experiment design stage. Work through the lessons in order when new to the area, o9
Open topic overview →
K-FoldK-Fold is a practical concept within Cross-Validation Methods. It helps turn the broader workflow stage “6 · Splitting, Validation & ExperimStratified K-FoldStratified K-Fold is a practical concept within Cross-Validation Methods. It helps turn the broader workflow stage “6 · Splitting, ValidatioRepeated K-FoldRepeated K-Fold is a practical concept within Cross-Validation Methods. It helps turn the broader workflow stage “6 · Splitting, Validation Group K-FoldGroup K-Fold is a practical concept within Cross-Validation Methods. It helps turn the broader workflow stage “6 · Splitting, Validation & EStratified Group K-FoldStratified Group K-Fold is a practical concept within Cross-Validation Methods. It helps turn the broader workflow stage “6 · Splitting, ValLeave-One-OutLeave-One-Out is a practical concept within Cross-Validation Methods. It helps turn the broader workflow stage “6 · Splitting, Validation & TimeSeriesSplitTimeSeriesSplit is a practical concept within Cross-Validation Methods. It helps turn the broader workflow stage “6 · Splitting, Validation Rolling and expanding windowsRolling and expanding windows is a practical concept within Cross-Validation Methods. It helps turn the broader workflow stage “6 · SplittinNested cross-validationNested cross-validation is a practical concept within Cross-Validation Methods. It helps turn the broader workflow stage “6 · Splitting, Val
Resampling & Statistical ComparisonResampling & Statistical Comparison groups the core ideas a learner needs at the 6 · splitting, validation & experiment design stage. Work through the lessons in order when new to 6
Open topic overview →
Bootstrap intuitionBootstrap intuition is a practical concept within Resampling & Statistical Comparison. It helps turn the broader workflow stage “6 · SplittiBootstrap confidence intervalsBootstrap confidence intervals is a practical concept within Resampling & Statistical Comparison. It helps turn the broader workflow stage “Permutation testsPermutation tests is a practical concept within Resampling & Statistical Comparison. It helps turn the broader workflow stage “6 · SplittingPaired fold comparisonsPaired fold comparisons is a practical concept within Resampling & Statistical Comparison. It helps turn the broader workflow stage “6 · SplRepeated runs and uncertaintyRepeated runs and uncertainty is a practical concept within Resampling & Statistical Comparison. It helps turn the broader workflow stage “6Multiple comparison cautionMultiple comparison caution is a practical concept within Resampling & Statistical Comparison. It helps turn the broader workflow stage “6 ·
Leakage PreventionLeakage Prevention groups the core ideas a learner needs at the 6 · splitting, validation & experiment design stage. Work through the lessons in order when new to the area, or use 8
Open topic overview →
Target leakageTarget leakage is a practical concept within Leakage Prevention. It helps turn the broader workflow stage “6 · Splitting, Validation & ExperPreprocessing leakagePreprocessing leakage is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness deFeature-selection leakageFeature-selection leakage changes how raw variables are represented for analysis or modelling. The transformation should preserve the informGroup leakageGroup leakage is a practical concept within Leakage Prevention. It helps turn the broader workflow stage “6 · Splitting, Validation & ExperiTemporal leakageTemporal leakage is a practical concept within Leakage Prevention. It helps turn the broader workflow stage “6 · Splitting, Validation & ExpDuplicate leakageDuplicate leakage is a practical concept within Leakage Prevention. It helps turn the broader workflow stage “6 · Splitting, Validation & ExHyperparameter-selection leakageHyperparameter-selection leakage is a practical concept within Leakage Prevention. It helps turn the broader workflow stage “6 · Splitting, Pipeline design as leakage controlPipeline design as leakage control changes how raw variables are represented for analysis or modelling. The transformation should preserve t
087 · Modelling & Training8 topics · 50 lessonsWorkflow stage
Ensemble LearningEnsembles combine multiple learners to improve robustness or accuracy by exploiting diversity among their errors. In this topic is expanded into smaller lessons so that definitions5
Open topic overview →
BaggingFits learners independently on resampled data/features and averages or votes. Random Forest is the canonical example. The important practicaBoostingFits learners sequentially so later stages focus on residual/error structure left by earlier stages. The important practical question is notStackingTrains heterogeneous base models and a meta-model on out-of-fold base predictions to learn how to combine them. The important practical quesVoting and averagingSimple hard voting, soft voting and weighted averaging can be effective when individual models are complementary. The important practical quLeakage-safe stackingMeta-model training must use predictions generated out-of-fold; in-sample base predictions leak target fit into the stacker. The important p
Graph ML FoundationsGraph machine learning represents entities as nodes and relationships as edges. GNNs learn by propagating, weighting and transforming information over this connectivity. In this to5
Open topic overview →
Graph representationA graph contains nodes, edges and optional node/edge/global features. Direction, edge type and temporal structure may all matter. The importMessage passingA node receives messages from neighbours, aggregates them with a permutation-invariant operator, then updates its hidden representation. TheReceptive fieldsAfter one layer, a node sees one-hop neighbours; deeper layers expand reach but can cause over-smoothing or over-squashing. The important prArchitecture familiesGCN normalises aggregation; GraphSAGE samples/aggregates; GAT learns attention weights; GIN uses expressive sum aggregation; graph transformGraph tasksNode classification/regression, edge/link prediction and graph-level prediction require different splitting and readout strategies. The impo
Regularisation & Early StoppingRegularisation controls effective model complexity so a learner captures reproducible structure instead of memorising noise. In this topic is expanded into smaller lessons so that 5
Open topic overview →
L2 regularisationPenalises squared parameter magnitude, encouraging smaller distributed weights and smoother solutions. The important practical question is nL1 regularisationPenalises absolute parameter magnitude and can drive coefficients exactly to zero, creating sparse models. The important practical question Elastic NetCombines L1 and L2 to balance sparsity with stability among correlated features. The important practical question is not only how the techniDropout and augmentationDeep-learning regularisers inject stochastic perturbations or varied examples, discouraging reliance on fragile pathways. The important pracEarly stoppingMonitors validation performance and stops optimisation when additional training no longer generalises. The important practical question is n
Baseline ModelsBaseline Models groups the core ideas a learner needs at the 7 · modelling & training stage. Work through the lessons in order when new to the area, or use them independently as a 5
Open topic overview →
Dummy classifier and regressorDummy classifier and regressor is a practical concept within Baseline Models. It helps turn the broader workflow stage “7 · Modelling & TraiMean / median regression baselineMean / median regression baseline is a practical concept within Baseline Models. It helps turn the broader workflow stage “7 · Modelling & TMajority-class baselineMajority-class baseline is a practical concept within Baseline Models. It helps turn the broader workflow stage “7 · Modelling & Training” iSimple linear baselineSimple linear baseline represents a family or practice in model building. The central idea is to define what structure can be learned, how mWhy baselines prevent self-deceptionWhy baselines prevent self-deception is a practical concept within Baseline Models. It helps turn the broader workflow stage “7 · Modelling
Supervised Model FamiliesSupervised Model Families groups the core ideas a learner needs at the 7 · modelling & training stage. Work through the lessons in order when new to the area, or use them independe8
Open topic overview →
Linear modelsLinear models represents a family or practice in model building. The central idea is to define what structure can be learned, how model qualNearest-neighbour methodsNearest-neighbour methods represents a family or practice in model building. The central idea is to define what structure can be learned, hoDecision treesDecision trees represents a family or practice in model building. The central idea is to define what structure can be learned, how model quaRandom forests and baggingRandom forests and bagging represents a family or practice in model building. The central idea is to define what structure can be learned, hGradient boostingGradient boosting represents a family or practice in model building. The central idea is to define what structure can be learned, how model Support vector machinesSupport vector machines is a practical concept within Supervised Model Families. It helps turn the broader workflow stage “7 · Modelling & TProbabilistic / Bayesian modelsProbabilistic / Bayesian models represents a family or practice in model building. The central idea is to define what structure can be learnNeural networksNeural networks represents a family or practice in model building. The central idea is to define what structure can be learned, how model qu
Unsupervised Model FamiliesUnsupervised Model Families groups the core ideas a learner needs at the 7 · modelling & training stage. Work through the lessons in order when new to the area, or use them indepen6
Open topic overview →
Partition clusteringPartition clustering represents a family or practice in model building. The central idea is to define what structure can be learned, how modDensity clusteringDensity clustering represents a family or practice in model building. The central idea is to define what structure can be learned, how modelHierarchical clusteringHierarchical clustering represents a family or practice in model building. The central idea is to define what structure can be learned, how Mixture modelsMixture models represents a family or practice in model building. The central idea is to define what structure can be learned, how model quaDimensionality reductionDimensionality reduction is a practical concept within Unsupervised Model Families. It helps turn the broader workflow stage “7 · Modelling Anomaly detectionAnomaly detection is a practical concept within Unsupervised Model Families. It helps turn the broader workflow stage “7 · Modelling & Train
Training DynamicsTraining Dynamics groups the core ideas a learner needs at the 7 · modelling & training stage. Work through the lessons in order when new to the area, or use them independently as 8
Open topic overview →
Loss functionsLoss functions is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends oGradient descentGradient descent is a practical concept within Training Dynamics. It helps turn the broader workflow stage “7 · Modelling & Training” into aLearning rateLearning rate is a practical concept within Training Dynamics. It helps turn the broader workflow stage “7 · Modelling & Training” into an eBatch sizeBatch size belongs to the operational phase where an analytical result becomes a maintained system. Production quality requires the data conEpochs and iterationsEpochs and iterations is a practical concept within Training Dynamics. It helps turn the broader workflow stage “7 · Modelling & Training” iConvergenceConvergence is a practical concept within Training Dynamics. It helps turn the broader workflow stage “7 · Modelling & Training” into an expRegularisationRegularisation is a practical concept within Training Dynamics. It helps turn the broader workflow stage “7 · Modelling & Training” into an Early stoppingEarly stopping is a practical concept within Training Dynamics. It helps turn the broader workflow stage “7 · Modelling & Training” into an
Class Imbalance StrategiesClass Imbalance Strategies groups the core ideas a learner needs at the 7 · modelling & training stage. Work through the lessons in order when new to the area, or use them independ8
Open topic overview →
Diagnose imbalanceDiagnose imbalance is a practical concept within Class Imbalance Strategies. It helps turn the broader workflow stage “7 · Modelling & TrainStratified evaluationStratified evaluation is a practical concept within Class Imbalance Strategies. It helps turn the broader workflow stage “7 · Modelling & TrClass weightsClass weights is a practical concept within Class Imbalance Strategies. It helps turn the broader workflow stage “7 · Modelling & Training” Random under-samplingRandom under-sampling is part of study design: it determines which units enter the dataset and therefore which population the analysis can lRandom over-samplingRandom over-sampling is part of study design: it determines which units enter the dataset and therefore which population the analysis can leSMOTE intuitionSMOTE intuition is a practical concept within Class Imbalance Strategies. It helps turn the broader workflow stage “7 · Modelling & TrainingThreshold movingThreshold moving is a practical concept within Class Imbalance Strategies. It helps turn the broader workflow stage “7 · Modelling & TraininPrecision-recall focusPrecision-recall focus is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness d
098 · Hyperparameter Optimisation & Model Selection6 topics · 36 lessonsWorkflow stage
Hyperparameter OptimisationHyperparameter optimisation searches configuration space for choices such as learning rate, tree depth, regularisation, kernel parameters and architecture size. In this topic is ex5
Open topic overview →
Manual and grid searchManual search is useful for intuition. Grid search is systematic but wastes trials when only a few dimensions strongly matter. The importantRandom searchRandom search samples independent configurations and often explores important dimensions more effectively than a full Cartesian grid for theBayesian optimisation and TPESequential model-based methods use earlier trials to focus future samples on promising regions while retaining exploration. The important prSuccessive Halving and HyperbandMulti-fidelity methods begin with many candidates at a small resource budget and allocate more epochs/data to strong candidates, reducing waPractical search designChoose a meaningful search space, log every trial, use early stopping when valid, and evaluate the final selected configuration on untouched
Learning Curves, Bias & VarianceLearning curves and complexity curves help diagnose underfitting, overfitting and whether additional data are likely to help. In this topic is expanded into smaller lessons so that5
Open topic overview →
High biasTraining and validation performance are both poor and close together; increasing model capacity or improving features may help. The importanHigh varianceTraining performance is strong but validation performance is substantially worse; regularisation, simpler models or more representative dataLearning curvesPlot performance versus training-set size. A persistent validation gap suggests variance; both curves plateauing poorly suggests bias. The iComplexity curvesVary depth, regularisation strength, feature count or another capacity parameter and compare training versus validation performance. The impData qualityCurve shape can also reflect label noise, dataset shift or leakage, so diagnostics must be interpreted in context. The important practical q
Model Selection & ComparisonModel comparison asks whether performance differences are stable, meaningful and worth the complexity rather than simply selecting the highest single score. In this topic is expand5
Open topic overview →
Baselines firstCompare against simple baselines and current operational rules before complex models. The important practical question is not only how the tPaired evaluationUse the same folds and test cases for candidate models so differences are paired rather than confounded by different samples. The important VariabilityReport fold/seed distributions, confidence intervals or bootstrap uncertainty where appropriate. The important practical question is not onlMultiple objectivesCompare predictive quality with calibration, latency, memory, interpretability, fairness and maintenance cost. The important practical questFinal selectionFreeze the selected pipeline and evaluate once on untouched data or a prospective period before deployment. The important practical question
Search Spaces & BudgetsSearch Spaces & Budgets groups the core ideas a learner needs at the 8 · hyperparameter optimisation & model selection stage. Work through the lessons in order when new to the area7
Open topic overview →
Hyperparameter vs learned parameterHyperparameter vs learned parameter is a practical concept within Search Spaces & Budgets. It helps turn the broader workflow stage “8 · HypLinear vs logarithmic rangesLinear vs logarithmic ranges represents a family or practice in model building. The central idea is to define what structure can be learned,Conditional parametersConditional parameters is a practical concept within Search Spaces & Budgets. It helps turn the broader workflow stage “8 · Hyperparameter OSearch budgetSearch budget is a practical concept within Search Spaces & Budgets. It helps turn the broader workflow stage “8 · Hyperparameter OptimisatiCross-validation costCross-validation cost is a practical concept within Search Spaces & Budgets. It helps turn the broader workflow stage “8 · Hyperparameter OpParallelismParallelism is a practical concept within Search Spaces & Budgets. It helps turn the broader workflow stage “8 · Hyperparameter OptimisationReproducible tuningReproducible tuning is a practical concept within Search Spaces & Budgets. It helps turn the broader workflow stage “8 · Hyperparameter Opti
Optimisation AlgorithmsOptimisation Algorithms groups the core ideas a learner needs at the 8 · hyperparameter optimisation & model selection stage. Work through the lessons in order when new to the area7
Open topic overview →
Grid searchGrid search is a practical concept within Optimisation Algorithms. It helps turn the broader workflow stage “8 · Hyperparameter OptimisationRandom searchRandom search is a practical concept within Optimisation Algorithms. It helps turn the broader workflow stage “8 · Hyperparameter OptimisatiBayesian optimisationBayesian optimisation is a practical concept within Optimisation Algorithms. It helps turn the broader workflow stage “8 · Hyperparameter OpTPETPE is a practical concept within Optimisation Algorithms. It helps turn the broader workflow stage “8 · Hyperparameter Optimisation & ModelSuccessive halvingSuccessive halving is a practical concept within Optimisation Algorithms. It helps turn the broader workflow stage “8 · Hyperparameter OptimHyperbandHyperband is a practical concept within Optimisation Algorithms. It helps turn the broader workflow stage “8 · Hyperparameter Optimisation &Evolutionary search intuitionEvolutionary search intuition is a practical concept within Optimisation Algorithms. It helps turn the broader workflow stage “8 · Hyperpara
Model Selection PracticeModel Selection Practice groups the core ideas a learner needs at the 8 · hyperparameter optimisation & model selection stage. Work through the lessons in order when new to the are7
Open topic overview →
Choose a primary metricChoose a primary metric is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness Multi-metric evaluationMulti-metric evaluation is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness Complexity vs performanceComplexity vs performance is a practical concept within Model Selection Practice. It helps turn the broader workflow stage “8 · HyperparametLatency and memory constraintsLatency and memory constraints belongs to the operational phase where an analytical result becomes a maintained system. Production quality rStability across foldsStability across folds is a practical concept within Model Selection Practice. It helps turn the broader workflow stage “8 · Hyperparameter One-standard-error rule intuitionOne-standard-error rule intuition is a practical concept within Model Selection Practice. It helps turn the broader workflow stage “8 · HypeFinal refit and untouched test setFinal refit and untouched test set is a practical concept within Model Selection Practice. It helps turn the broader workflow stage “8 · Hyp
109 · Evaluation, Metrics & Diagnostics13 topics · 85 lessonsWorkflow stage
Calibration & Decision ThresholdsCalibration concerns whether predicted probabilities match observed frequencies; thresholding converts scores/probabilities into operational decisions. In this topic is expanded in5
Open topic overview →
Threshold selectionChoose thresholds using the decision objective: sensitivity target, precision target, expected cost, workload or utility. The important pracReliability diagramsBin predicted probabilities and compare mean confidence with observed outcome frequency. Deviations from the diagonal indicate miscalibratioPlatt / sigmoid scalingFits a logistic mapping from raw scores to probabilities on calibration data. The important practical question is not only how the techniqueIsotonic regressionLearns a flexible monotonic calibration mapping; it can fit complex miscalibration but needs more calibration data. The important practical Temperature scalingCommon in neural networks: a single temperature rescales logits, preserving class ranking while changing confidence. The important practical
Classification MetricsClassification metrics answer different questions: how often predictions are correct, how well minority positives are found, how reliable probabilities are, and how ranking quality5
Open topic overview →
Confusion matrixTP, TN, FP and FN are the building blocks of threshold-based metrics. Always inspect the counts as well as derived scores. The important praPrecision, recall and F-scoresPrecision measures purity of positive predictions; recall/sensitivity measures positive coverage. F1 balances them, while Fβ can emphasise rROC-AUC and PR-AUCROC-AUC evaluates ranking across thresholds; PR-AUC is often more informative when positives are rare because it focuses on positive predictProbability metricsLog loss and Brier score evaluate probabilistic confidence, penalising confident wrong predictions. Calibration should be inspected when proBalanced and correlation metricsBalanced accuracy averages class recalls. MCC and Cohen’s κ can offer robust single-number summaries in imbalanced settings, but should stil
Clustering Validation MetricsClustering evaluation measures compactness, separation or agreement with known labels while recognising that no single score defines a scientifically meaningful cluster. In this to5
Open topic overview →
Silhouette coefficientCompares within-cluster cohesion with separation from the nearest alternative cluster. Higher is generally better, but convex distance-basedDavies–Bouldin indexMeasures average similarity between each cluster and its most similar alternative. Lower values indicate more separated compact clusters. ThCalinski–Harabasz indexCompares between-cluster dispersion with within-cluster dispersion and often favours well-separated spherical structure. The important practExternal validationAdjusted Rand Index and Normalised Mutual Information compare clusters with known labels while correcting or normalising agreement. The impoStabilityRepeat clustering under resampling, initialisation or perturbation. Stable structure often matters more than marginal changes in one interna
Predictive UncertaintyUncertainty estimation distinguishes what the model predicts from how confident it should be, including noise intrinsic to data and uncertainty due to limited model knowledge. In t5
Open topic overview →
Aleatoric uncertaintyIrreducible variability or measurement noise in the data-generating process. The important practical question is not only how the technique Epistemic uncertaintyUncertainty about model parameters/functions due to limited evidence; it may reduce with informative new data. The important practical questPredictive intervalsRegression intervals communicate a range of plausible outcomes; coverage should be validated empirically. The important practical question iBayesian and ensemble approachesGaussian Processes, Bayesian neural methods, deep ensembles and MC dropout approximate distributions over predictions. The important practicConformal predictionUses calibration residuals/scores to create sets or intervals with finite-sample coverage guarantees under exchangeability assumptions. The
Regression Metrics & ResidualsRegression metrics quantify prediction error in different units and with different sensitivity to large mistakes. Residual diagnostics reveal patterns that aggregate metrics hide. 5
Open topic overview →
MAE and median absolute errorMAE is interpretable in target units and treats each absolute error linearly. Median absolute error is robust to extreme residuals. The impoMSE and RMSESquared error emphasises large misses. RMSE returns to target units and is useful when large errors are disproportionately costly. The imporR² and adjusted R²R² compares residual variance with a mean baseline. A good R² does not guarantee unbiased or well-calibrated predictions; adjusted R² is maiPercentage and log errorsMAPE is intuitive but problematic near zero. sMAPE and log-scale losses may help for multiplicative targets but change the optimisation meanResidual analysisPlot residuals against predictions, key features and time. Structure, changing variance or systematic bias indicates missing relationships o
Classification Metrics: CoreClassification Metrics: Core groups the core ideas a learner needs at the 9 · evaluation, metrics & diagnostics stage. Work through the lessons in order when new to the area, or us8
Open topic overview →
Confusion matrixConfusion matrix is a practical concept within Classification Metrics: Core. It helps turn the broader workflow stage “9 · Evaluation, MetriAccuracyAccuracy is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends on whetBalanced accuracyBalanced accuracy is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness dependPrecisionPrecision is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends on wheRecall / sensitivityRecall / sensitivity is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depSpecificitySpecificity is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends on wF1 scoreF1 score is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends on whetF-beta scoreF-beta score is a practical concept within Classification Metrics: Core. It helps turn the broader workflow stage “9 · Evaluation, Metrics &
Classification Metrics: Ranking & ProbabilityClassification Metrics: Ranking & Probability groups the core ideas a learner needs at the 9 · evaluation, metrics & diagnostics stage. Work through the lessons in order when new t8
Open topic overview →
ROC curveROC curve is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends on wheROC-AUCROC-AUC is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends on whethPrecision-recall curvePrecision-recall curve is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness dAverage precision / PR-AUCAverage precision / PR-AUC is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulneLog lossLog loss is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends on whetBrier scoreBrier score is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends on wTop-k accuracyTop-k accuracy is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends oMCC and Cohen’s kappaMCC and Cohen’s kappa is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness de
Regression MetricsRegression Metrics groups the core ideas a learner needs at the 9 · evaluation, metrics & diagnostics stage. Work through the lessons in order when new to the area, or use them ind10
Open topic overview →
MAEMAE is a practical concept within Regression Metrics. It helps turn the broader workflow stage “9 · Evaluation, Metrics & Diagnostics” into MSEMSE is a practical concept within Regression Metrics. It helps turn the broader workflow stage “9 · Evaluation, Metrics & Diagnostics” into RMSERMSE is a practical concept within Regression Metrics. It helps turn the broader workflow stage “9 · Evaluation, Metrics & Diagnostics” intoR-squaredR-squared is a practical concept within Regression Metrics. It helps turn the broader workflow stage “9 · Evaluation, Metrics & Diagnostics”Adjusted R-squaredAdjusted R-squared is a practical concept within Regression Metrics. It helps turn the broader workflow stage “9 · Evaluation, Metrics & DiaMAPEMAPE is a practical concept within Regression Metrics. It helps turn the broader workflow stage “9 · Evaluation, Metrics & Diagnostics” intosMAPEsMAPE is a practical concept within Regression Metrics. It helps turn the broader workflow stage “9 · Evaluation, Metrics & Diagnostics” intMSLE and RMSLEMSLE and RMSLE is a practical concept within Regression Metrics. It helps turn the broader workflow stage “9 · Evaluation, Metrics & DiagnosMedian absolute errorMedian absolute error is a practical concept within Regression Metrics. It helps turn the broader workflow stage “9 · Evaluation, Metrics & Huber lossHuber loss is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends on wh
Regression DiagnosticsRegression Diagnostics groups the core ideas a learner needs at the 9 · evaluation, metrics & diagnostics stage. Work through the lessons in order when new to the area, or use them7
Open topic overview →
Residual vs fitted plotResidual vs fitted plot is a communication and diagnostic technique that maps data or results into a visual form. A useful visual makes compResidual distributionResidual distribution is a practical concept within Regression Diagnostics. It helps turn the broader workflow stage “9 · Evaluation, MetricHeteroscedasticity intuitionHeteroscedasticity intuition is a practical concept within Regression Diagnostics. It helps turn the broader workflow stage “9 · Evaluation,Systematic biasSystematic bias is a practical concept within Regression Diagnostics. It helps turn the broader workflow stage “9 · Evaluation, Metrics & DiInfluential observationsInfluential observations is a practical concept within Regression Diagnostics. It helps turn the broader workflow stage “9 · Evaluation, MetPrediction interval coveragePrediction interval coverage is a practical concept within Regression Diagnostics. It helps turn the broader workflow stage “9 · Evaluation,Segment-wise error analysisSegment-wise error analysis is a practical concept within Regression Diagnostics. It helps turn the broader workflow stage “9 · Evaluation,
Clustering MetricsClustering Metrics groups the core ideas a learner needs at the 9 · evaluation, metrics & diagnostics stage. Work through the lessons in order when new to the area, or use them ind6
Open topic overview →
Silhouette scoreSilhouette score is a practical concept within Clustering Metrics. It helps turn the broader workflow stage “9 · Evaluation, Metrics & DiagnDavies-Bouldin indexDavies-Bouldin index is a practical concept within Clustering Metrics. It helps turn the broader workflow stage “9 · Evaluation, Metrics & DCalinski-Harabasz indexCalinski-Harabasz index is a practical concept within Clustering Metrics. It helps turn the broader workflow stage “9 · Evaluation, Metrics Adjusted Rand IndexAdjusted Rand Index is a practical concept within Clustering Metrics. It helps turn the broader workflow stage “9 · Evaluation, Metrics & DiNormalized mutual informationNormalized mutual information is a practical concept within Clustering Metrics. It helps turn the broader workflow stage “9 · Evaluation, MeCluster stabilityCluster stability is a practical concept within Clustering Metrics. It helps turn the broader workflow stage “9 · Evaluation, Metrics & Diag
Forecasting MetricsForecasting Metrics groups the core ideas a learner needs at the 9 · evaluation, metrics & diagnostics stage. Work through the lessons in order when new to the area, or use them in6
Open topic overview →
Forecast MAE and RMSEForecast MAE and RMSE is a practical concept within Forecasting Metrics. It helps turn the broader workflow stage “9 · Evaluation, Metrics &MAPE limitationsMAPE limitations is a practical concept within Forecasting Metrics. It helps turn the broader workflow stage “9 · Evaluation, Metrics & DiagsMAPEsMAPE is a practical concept within Forecasting Metrics. It helps turn the broader workflow stage “9 · Evaluation, Metrics & Diagnostics” inMASEMASE is a practical concept within Forecasting Metrics. It helps turn the broader workflow stage “9 · Evaluation, Metrics & Diagnostics” intPinball lossPinball loss is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usefulness depends on Backtesting metrics by horizonBacktesting metrics by horizon is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its usef
Calibration, Thresholds & Decision CostsCalibration, Thresholds & Decision Costs groups the core ideas a learner needs at the 9 · evaluation, metrics & diagnostics stage. Work through the lessons in order when new to the8
Open topic overview →
Probability calibrationProbability calibration is a practical concept within Calibration, Thresholds & Decision Costs. It helps turn the broader workflow stage “9 Reliability diagramsReliability diagrams is a practical concept within Calibration, Thresholds & Decision Costs. It helps turn the broader workflow stage “9 · EPlatt scalingPlatt scaling is a practical concept within Calibration, Thresholds & Decision Costs. It helps turn the broader workflow stage “9 · EvaluatiIsotonic calibrationIsotonic calibration is a practical concept within Calibration, Thresholds & Decision Costs. It helps turn the broader workflow stage “9 · ETemperature scalingTemperature scaling is a practical concept within Calibration, Thresholds & Decision Costs. It helps turn the broader workflow stage “9 · EvThreshold optimisationThreshold optimisation is a practical concept within Calibration, Thresholds & Decision Costs. It helps turn the broader workflow stage “9 ·Cost-sensitive thresholdsCost-sensitive thresholds is a practical concept within Calibration, Thresholds & Decision Costs. It helps turn the broader workflow stage “Abstention / reject optionAbstention / reject option is a practical concept within Calibration, Thresholds & Decision Costs. It helps turn the broader workflow stage
Uncertainty & IntervalsUncertainty & Intervals groups the core ideas a learner needs at the 9 · evaluation, metrics & diagnostics stage. Work through the lessons in order when new to the area, or use the7
Open topic overview →
Confidence vs prediction intervalsConfidence vs prediction intervals is a practical concept within Uncertainty & Intervals. It helps turn the broader workflow stage “9 · EvalAleatoric uncertaintyAleatoric uncertainty is a practical concept within Uncertainty & Intervals. It helps turn the broader workflow stage “9 · Evaluation, MetriEpistemic uncertaintyEpistemic uncertainty is a practical concept within Uncertainty & Intervals. It helps turn the broader workflow stage “9 · Evaluation, MetriBootstrap intervalsBootstrap intervals is a practical concept within Uncertainty & Intervals. It helps turn the broader workflow stage “9 · Evaluation, MetricsBayesian predictive uncertaintyBayesian predictive uncertainty is a practical concept within Uncertainty & Intervals. It helps turn the broader workflow stage “9 · EvaluatConformal predictionConformal prediction is a practical concept within Uncertainty & Intervals. It helps turn the broader workflow stage “9 · Evaluation, MetricCoverage and interval widthCoverage and interval width is a practical concept within Uncertainty & Intervals. It helps turn the broader workflow stage “9 · Evaluation,
1110 · Interpretation & Explainability4 topics · 23 lessonsWorkflow stage
Explainability & Model InterpretationExplainability techniques describe global model behaviour or why a specific prediction changed, but explanations are approximations whose assumptions must be understood. In this to5
Open topic overview →
Global importancePermutation importance measures performance loss when a feature is disrupted; tree impurity importance is fast but can be biased. The importPDP and ICEPartial Dependence averages predictions across a feature grid; ICE shows individual trajectories and exposes heterogeneous effects. The impoSHAP conceptsShapley-based attributions distribute prediction difference among features using a game-theoretic framework; background/reference choices maLocal surrogate methodsLIME approximates the model locally around one instance; fidelity and perturbation strategy should be checked. The important practical questDeep visual explanationsIntegrated Gradients, Grad-CAM and attention visualisations expose gradients/activation patterns, but saliency is not identical to causal re
Interpretable ModelsInterpretable Models groups the core ideas a learner needs at the 10 · interpretation & explainability stage. Work through the lessons in order when new to the area, or use them in5
Open topic overview →
Linear coefficientsLinear coefficients represents a family or practice in model building. The central idea is to define what structure can be learned, how modeOdds ratiosOdds ratios is a practical concept within Interpretable Models. It helps turn the broader workflow stage “10 · Interpretation & ExplainabiliTree rulesTree rules represents a family or practice in model building. The central idea is to define what structure can be learned, how model qualityMonotonic relationshipsMonotonic relationships is a practical concept within Interpretable Models. It helps turn the broader workflow stage “10 · Interpretation & Global vs local explanationsGlobal vs local explanations is a practical concept within Interpretable Models. It helps turn the broader workflow stage “10 · Interpretati
Model-Agnostic ExplainabilityModel-Agnostic Explainability groups the core ideas a learner needs at the 10 · interpretation & explainability stage. Work through the lessons in order when new to the area, or us7
Open topic overview →
Permutation importancePermutation importance is a practical concept within Model-Agnostic Explainability. It helps turn the broader workflow stage “10 · InterpretPartial dependencePartial dependence is a practical concept within Model-Agnostic Explainability. It helps turn the broader workflow stage “10 · InterpretatioICE plotsICE plots is a communication and diagnostic technique that maps data or results into a visual form. A useful visual makes comparisons easy, SHAP intuitionSHAP intuition is a practical concept within Model-Agnostic Explainability. It helps turn the broader workflow stage “10 · Interpretation & LIME intuitionLIME intuition is a practical concept within Model-Agnostic Explainability. It helps turn the broader workflow stage “10 · Interpretation & Counterfactual explanationsCounterfactual explanations is a practical concept within Model-Agnostic Explainability. It helps turn the broader workflow stage “10 · InteCorrelated-feature caveatsCorrelated-feature caveats changes how raw variables are represented for analysis or modelling. The transformation should preserve the infor
Deep Learning ExplainabilityDeep Learning Explainability groups the core ideas a learner needs at the 10 · interpretation & explainability stage. Work through the lessons in order when new to the area, or use6
Open topic overview →
Saliency mapsSaliency maps is a practical concept within Deep Learning Explainability. It helps turn the broader workflow stage “10 · Interpretation & ExIntegrated gradientsIntegrated gradients is a practical concept within Deep Learning Explainability. It helps turn the broader workflow stage “10 · InterpretatiGrad-CAMGrad-CAM is a practical concept within Deep Learning Explainability. It helps turn the broader workflow stage “10 · Interpretation & ExplainAttention visualisationAttention visualisation is a communication and diagnostic technique that maps data or results into a visual form. A useful visual makes compEmbedding visualisationEmbedding visualisation is a communication and diagnostic technique that maps data or results into a visual form. A useful visual makes compExplanation stabilityExplanation stability is a practical concept within Deep Learning Explainability. It helps turn the broader workflow stage “10 · Interpretat
1211 · Post-processing, Reporting & Communication4 topics · 24 lessonsWorkflow stage
Prediction Post-ProcessingPost-processing transforms raw model outputs into final usable predictions while respecting domain constraints, calibration and operational decisions. In this topic is expanded int5
Open topic overview →
ClassificationThreshold optimisation, probability calibration, top-k rules, class-specific thresholds and abstention/reject options. The important practicRegressionInverse target transforms, bias correction, constraint clipping and predictive intervals. The important practical question is not only how tForecastingTrend/seasonality restoration, smoothing, reconciliation across hierarchies and domain constraints. The important practical question is not Computer visionConfidence filtering, Intersection-over-Union thresholds and Non-Maximum Suppression turn dense detector outputs into final boxes. The imporGenerative/NLPTemperature, top-k/top-p sampling, beam search and repetition controls shape decoded outputs after the model produces logits. The important
Results, Reporting & Decision CommunicationTranslate analysis and model outputs into clear evidence, uncertainty, limitations and actionable decisions for technical and non-technical audiences. In this topic is expanded int5
Open topic overview →
Choose the message and audienceDifferent audiences need different levels of method detail, uncertainty and operational context. The important practical question is not onlTables and visualisationGood visuals expose patterns and uncertainty without decorative clutter. The important practical question is not only how the technique is dUncertainty and limitationsReport variability, confidence/prediction intervals, data limitations and plausible failure modes. The important practical question is not oModel cards and experiment summariesStructured documentation captures intended use, data, metrics, subgroups, limitations and reproducibility details. The important practical qDecision thresholds and action plansOutputs become useful when mapped to actions, owners and escalation rules. The important practical question is not only how the technique is
Reporting & Visual StorytellingReporting & Visual Storytelling groups the core ideas a learner needs at the 11 · post-processing, reporting & communication stage. Work through the lessons in order when new to th7
Open topic overview →
Audience and decision messageAudience and decision message is a practical concept within Reporting & Visual Storytelling. It helps turn the broader workflow stage “11 · Choose the right tableChoose the right table is a communication and diagnostic technique that maps data or results into a visual form. A useful visual makes compaChoose the right chartChoose the right chart is a communication and diagnostic technique that maps data or results into a visual form. A useful visual makes compaAnnotate uncertaintyAnnotate uncertainty is a practical concept within Reporting & Visual Storytelling. It helps turn the broader workflow stage “11 · Post-procAvoid chartjunkAvoid chartjunk is a communication and diagnostic technique that maps data or results into a visual form. A useful visual makes comparisons Executive summaryExecutive summary is a practical concept within Reporting & Visual Storytelling. It helps turn the broader workflow stage “11 · Post-processTechnical appendixTechnical appendix is a practical concept within Reporting & Visual Storytelling. It helps turn the broader workflow stage “11 · Post-proces
Reproducibility & DocumentationReproducibility & Documentation groups the core ideas a learner needs at the 11 · post-processing, reporting & communication stage. Work through the lessons in order when new to th7
Open topic overview →
Experiment trackingExperiment tracking is a practical concept within Reproducibility & Documentation. It helps turn the broader workflow stage “11 · Post-proceData versioningData versioning is a practical concept within Reproducibility & Documentation. It helps turn the broader workflow stage “11 · Post-processinModel cardsModel cards represents a family or practice in model building. The central idea is to define what structure can be learned, how model qualitDatasheets for datasetsDatasheets for datasets is a practical concept within Reproducibility & Documentation. It helps turn the broader workflow stage “11 · Post-pEnvironment captureEnvironment capture is a practical concept within Reproducibility & Documentation. It helps turn the broader workflow stage “11 · Post-proceRandom seedsRandom seeds is a practical concept within Reproducibility & Documentation. It helps turn the broader workflow stage “11 · Post-processing, Decision logsDecision logs is a practical concept within Reproducibility & Documentation. It helps turn the broader workflow stage “11 · Post-processing,
1312 · Deployment, Monitoring & Improvement4 topics · 27 lessonsWorkflow stage
Deployment, Drift & MonitoringMonitoring checks whether data, predictions, performance and operational assumptions remain valid after deployment. In this topic is expanded into smaller lessons so that definitio5
Open topic overview →
Data driftFeature distributions change relative to the development baseline. Drift is a warning signal, not proof of performance degradation. The impoConcept driftThe relationship between inputs and outcomes changes, so historical decision boundaries or regression functions become stale. The important Performance monitoringWhen labels arrive, track task metrics by time and important segments. Delayed labels require proxy and drift signals in the meantime. The iCalibration and threshold monitoringProbability calibration and optimal decision thresholds can shift even when ranking remains stable. The important practical question is not Retraining governanceDefine triggers, approval, rollback, versioning and reproducible evaluation before automated retraining. The important practical question is
Deployment PatternsDeployment Patterns groups the core ideas a learner needs at the 12 · deployment, monitoring & improvement stage. Work through the lessons in order when new to the area, or use the7
Open topic overview →
Batch scoringBatch scoring belongs to the operational phase where an analytical result becomes a maintained system. Production quality requires the data Online API inferenceOnline API inference belongs to the operational phase where an analytical result becomes a maintained system. Production quality requires thStreaming inferenceStreaming inference belongs to the operational phase where an analytical result becomes a maintained system. Production quality requires theEdge deployment intuitionEdge deployment intuition belongs to the operational phase where an analytical result becomes a maintained system. Production quality requirModel serialisationModel serialisation represents a family or practice in model building. The central idea is to define what structure can be learned, how modeServing the preprocessing pipelineServing the preprocessing pipeline is an evaluation quantity that compresses a particular aspect of predictive behaviour into a number. Its Latency and throughput basicsLatency and throughput basics belongs to the operational phase where an analytical result becomes a maintained system. Production quality re
Production MonitoringProduction Monitoring groups the core ideas a learner needs at the 12 · deployment, monitoring & improvement stage. Work through the lessons in order when new to the area, or use t8
Open topic overview →
Input schema checksInput schema checks is a practical concept within Production Monitoring. It helps turn the broader workflow stage “12 · Deployment, MonitoriData-quality monitoringData-quality monitoring belongs to the operational phase where an analytical result becomes a maintained system. Production quality requiresData driftData drift is a practical concept within Production Monitoring. It helps turn the broader workflow stage “12 · Deployment, Monitoring & ImprConcept driftConcept drift is a practical concept within Production Monitoring. It helps turn the broader workflow stage “12 · Deployment, Monitoring & IPerformance monitoringPerformance monitoring belongs to the operational phase where an analytical result becomes a maintained system. Production quality requires Calibration monitoringCalibration monitoring belongs to the operational phase where an analytical result becomes a maintained system. Production quality requires Fairness monitoringFairness monitoring belongs to the operational phase where an analytical result becomes a maintained system. Production quality requires theLatency and failure monitoringLatency and failure monitoring belongs to the operational phase where an analytical result becomes a maintained system. Production quality r
Lifecycle & RetrainingLifecycle & Retraining groups the core ideas a learner needs at the 12 · deployment, monitoring & improvement stage. Work through the lessons in order when new to the area, or use 7
Open topic overview →
Retraining triggersRetraining triggers belongs to the operational phase where an analytical result becomes a maintained system. Production quality requires theScheduled vs event-driven retrainingScheduled vs event-driven retraining belongs to the operational phase where an analytical result becomes a maintained system. Production quaShadow evaluationShadow evaluation is a practical concept within Lifecycle & Retraining. It helps turn the broader workflow stage “12 · Deployment, MonitorinChampion-challenger testingChampion-challenger testing is a practical concept within Lifecycle & Retraining. It helps turn the broader workflow stage “12 · Deployment,RollbackRollback belongs to the operational phase where an analytical result becomes a maintained system. Production quality requires the data contrApproval and governanceApproval and governance is a practical concept within Lifecycle & Retraining. It helps turn the broader workflow stage “12 · Deployment, MonPost-deployment feedback loopsPost-deployment feedback loops belongs to the operational phase where an analytical result becomes a maintained system. Production quality r