Visualisation & Storytelling · Lesson 64

Scatterplots

Scatterplots is an exploratory data-analysis technique.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Scatterplots actually means

Scatterplots is an exploratory data-analysis technique. EDA is structured investigation of distributions, groups, relationships, missingness and anomalies before stronger inferential or predictive claims are made.

Scatterplots matters because visual encodings determine what comparisons a reader can perceive quickly and accurately. Good storytelling does not decorate evidence; it selects and labels the representation that best supports the analytical question.

Deeper walkthrough

Read Scatterplots as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Start with the unit of analysis and data-quality profile. Stage 2: Use univariate summaries to inspect centre, spread, shape and unusual values. Stage 3: Compare groups using both counts and distributions rather than only averages. Final checkpoint: Record observations as hypotheses or data-quality findings, not as automatic causal claims.

Mechanism

Follow the transformation

Start with the unit of analysis and data-quality profile.

Use univariate summaries to inspect centre, spread, shape and unusual values.

Compare groups using both counts and distributions rather than only averages.

Evidence

Know what would convince you

  • Reconcile plotted marks/bars/bins/lines with a few source values and the intended aggregation.
  • Check axis limits, units, category order and any binning/smoothing choices explicitly.
Useful distinctionUnivariate: One variable: distribution, counts, centre and spread.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Start with the unit of analysis…

Start with the unit of analysis and data-quality profile. For Scatterplots, identify the exact state before this stage, the operation or rule applied here, and the observable state afterwards so the mechanism remains inspectable.

State focus: identify exactly what changed at this stage and what observable evidence confirms that change.
How it works

Trace the mechanism step by step

  1. Start with the unit of analysis and data-quality profile.
  2. Use univariate summaries to inspect centre, spread, shape and unusual values.
  3. Compare groups using both counts and distributions rather than only averages.
  4. Inspect bivariate relationships with plots and appropriate association measures.
  5. Look for missingness, clusters, nonlinear structure, time effects and subgroup patterns that could change a conclusion.
  6. Record observations as hypotheses or data-quality findings, not as automatic causal claims.
Worked demonstration

Make the concept concrete

Demonstration

Python example

# Step 1 — Compute the right-hand expression and store its result in `x` for the next step.
x = [1,2,3,4,5]
# Step 2 — Compute the right-hand expression and store its result in `y` for the next step.
y = [2,4,5,8,9]
# Step 3 — Iterate through the collection so the indented block is applied once for each item.
for a,b in zip(x,y): print(f"({a},{b})")
Expected / illustrative result
The paired points rise overall. A scatter plot preserves each pair, unlike two separate histograms.
Interpret the result.

For Scatterplots, identify which source values create each important mark/position, then check whether scale, ordering, aggregation or binning could change the visual conclusion.

Distinctions & related ideas

Know what this is — and what it is not

UnivariateOne variable: distribution, counts, centre and spread.
BivariateTwo variables: association or group differences.
MultivariateSeveral variables: interactions, confounding, conditional patterns.
Confirmatory analysisTests or models a pre-specified claim; should be distinguished from open-ended exploration.
Use deliberately

When it is appropriate

Use Scatterplots when its visual encoding matches the variable types and the comparison/pattern the reader needs to see.

Boundary conditions

When to stop or reconsider

Choose a different chart or representation when this encoding hides distribution, order, uncertainty or observation-level structure, or when overplotting/scale choices would make the display misleading.

Common mistakes

Failure modes to recognise

  • Using a chart type whose marks/axes do not match the data types or analytical question.
  • Allowing scale, binning, category order or aggregation to create a visual impression the raw values do not support.
  • Adding colour, labels or decoration without a clear encoding purpose or accessible alternative.
Verification

How to check the result

  • Reconcile plotted marks/bars/bins/lines with a few source values and the intended aggregation.
  • Check axis limits, units, category order and any binning/smoothing choices explicitly.
  • Change one data value or filtering rule and predict which visual element should move before regenerating the figure.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Scatterplots. First start with the unit of analysis and data-quality profile. Then use univariate summaries to inspect centre, spread, shape and unusual values. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Use a handful of values and label the axes/units. Point to each visual mark and identify the source value or aggregation that created it.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Scatterplots, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Start with the unit of analysis and data-quality profile.
Step 2Use univariate summaries to inspect centre, spread, shape and unusual values.
Step 3Compare groups using both counts and distributions rather than only averages.
Step 4Inspect bivariate relationships with plots and appropriate association measures.
Lesson summary

What to remember

  • Scatterplots is an exploratory data-analysis technique. EDA is structured investigation of distributions, groups, relationships, missingness and anomalies before stronger inferential or predictive claims are made.
  • Start with the unit of analysis and data-quality profile.
  • Using a chart type whose marks/axes do not match the data types or analytical question.
  • Reconcile plotted marks/bars/bins/lines with a few source values and the intended aggregation.