Exploratory Data Analysis · Lesson 56

Boxplots and Outliers

Boxplots and Outliers is a data-quality decision, not merely a cleaning command.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Boxplots and Outliers actually means

Boxplots and Outliers is a data-quality decision, not merely a cleaning command. The correct treatment depends on how the issue arose, whether it carries information, and how the treatment changes the population or downstream model.

Boxplots and Outliers matters because exploratory analysis is where structure, anomalies and plausible relationships become visible before stronger claims are made. The goal is to generate and test questions while preserving uncertainty and data-quality context.

Deeper walkthrough

Read Boxplots and Outliers as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Measure the problem by column, row group and time/segment. Stage 2: Investigate the data-generating process before choosing a fix. Stage 3: Choose deletion, correction, imputation, transformation or retention with an explicit reason. Final checkpoint: Record an indicator or audit trail when the fact that a value was missing/changed may itself matter.

Mechanism

Follow the transformation

Measure the problem by column, row group and time/segment.

Investigate the data-generating process before choosing a fix.

Choose deletion, correction, imputation, transformation or retention with an explicit reason.

Evidence

Know what would convince you

  • Compare row/column counts, dtypes and missing values before and after the operation.
  • Trace a few representative rows or one group manually from source values to result.
Useful distinctionDeletion: Removes affected rows/columns; can bias the population.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Measure the problem by column

Measure the problem by column, row group and time/segment. For Boxplots and Outliers, identify the exact state before this stage, the operation or rule applied here, and the observable state afterwards so the mechanism remains inspectable.

State focus: identify exactly what changed at this stage and what observable evidence confirms that change.
How it works

Trace the mechanism step by step

  1. Measure the problem by column, row group and time/segment.
  2. Investigate the data-generating process before choosing a fix.
  3. Choose deletion, correction, imputation, transformation or retention with an explicit reason.
  4. Fit learned preprocessing only on training data when predictive modelling is involved.
  5. Record an indicator or audit trail when the fact that a value was missing/changed may itself matter.
Worked demonstration

Make the concept concrete

Demonstration

Python / NumPy example

# Step 1 — Import the module so its functions/classes are available to the rest of this example.
import numpy as np
# Step 2 — Construct `x` as an array so vectorised numerical operations can be applied consistently.
x = np.array([10,11,11,12,12,13,40])
# Step 3 — Compute the right-hand expression and store its result in `q1,q3` for the next step.
q1,q3 = np.quantile(x,[.25,.75]); iqr=q3-q1
# Step 4 — Compute the right-hand expression and store its result in `lo,hi` for the next step.
lo,hi=q1-1.5*iqr,q3+1.5*iqr
# Step 5 — Display the current value explicitly so the result/state can be inspected during execution.
print("fences:",lo,hi)
# Step 6 — Display the current value explicitly so the result/state can be inspected during execution.
print("flagged:",x[(x<lo)|(x>hi)])
Expected / illustrative result
40 is flagged by the 1.5×IQR rule. That is a review flag, not automatic evidence the value is wrong.
Interpret the result.

For Boxplots and Outliers, trace representative source rows/columns into the result and reconcile row counts, dtypes, keys or missing values that the operation could change.

Distinctions & related ideas

Know what this is — and what it is not

DeletionRemoves affected rows/columns; can bias the population.
Simple imputationUses a fixed statistic/category; easy but shrinks variability.
Model-based imputationUses relationships with other variables; more assumptions and leakage risk.
Missing indicatorPreserves information about whether a value was missing.
Use deliberately

When it is appropriate

Use Boxplots and Outliers when the data are naturally tabular and row grain, column meaning, keys and dtypes can be stated explicitly.

Boundary conditions

When to stop or reconsider

Reconsider the operation if row identity/grain is unclear, join keys are not validated, chained transformations hide state, or the task is better expressed with a simpler table operation.

Common mistakes

Failure modes to recognise

  • Changing row grain or row count without noticing it.
  • Joining/grouping on keys whose uniqueness or missingness was never checked.
  • Interpreting a derived column or aggregation without reconciling it to source rows and units.
Verification

How to check the result

  • Compare row/column counts, dtypes and missing values before and after the operation.
  • Trace a few representative rows or one group manually from source values to result.
  • For joins/reshapes/grouping, verify key uniqueness/cardinality and reconcile totals where totals should be preserved.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Boxplots and Outliers. First measure the problem by column, row group and time/segment. Then investigate the data-generating process before choosing a fix. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Work with 4–8 rows that contain the exact key/category/missing-value pattern you want to understand. Trace one row or group all the way through.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Boxplots and Outliers, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Measure the problem by column, row group and time/segment.
Step 2Investigate the data-generating process before choosing a fix.
Step 3Choose deletion, correction, imputation, transformation or retention with an explicit reason.
Step 4Fit learned preprocessing only on training data when predictive modelling is involved.
Lesson summary

What to remember

  • Boxplots and Outliers is a data-quality decision, not merely a cleaning command. The correct treatment depends on how the issue arose, whether it carries information, and how the treatment changes the population or downstream model.
  • Measure the problem by column, row group and time/segment.
  • Changing row grain or row count without noticing it.
  • Compare row/column counts, dtypes and missing values before and after the operation.