Statistical Reasoning · Lesson 81

Correlation vs Causation

Correlation summarises the direction and strength of association between variables.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Correlation vs Causation actually means

Correlation summarises the direction and strength of association between variables. Pearson correlation measures linear association on the original numeric scale; Spearman correlation measures monotonic association after converting values to ranks. A coefficient is not a causal estimate and should be interpreted together with the scatter plot and data-generating context.

Correlation vs Causation matters because sample summaries vary even when the underlying process has not changed. Statistical reasoning provides a language for variation and uncertainty so analysts do not treat every observed difference as a reliable population difference.

Deeper walkthrough

Read Correlation vs Causation as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Pair observations correctly and handle missing pairs deliberately. Stage 2: Plot the variables to check shape, clusters and outliers. Stage 3: Compute Pearson r for linear association or Spearman rho for monotonic ranked association when appropriate. Final checkpoint: Never turn correlation into a causal statement without a design that identifies causality.

Mechanism

Follow the transformation

Pair observations correctly and handle missing pairs deliberately.

Plot the variables to check shape, clusters and outliers.

Compute Pearson r for linear association or Spearman rho for monotonic ranked association when appropriate.

Evidence

Know what would convince you

  • Compute a tiny example by hand or simulate a simple case where the expected behaviour is known.
  • Check units, sample size, denominator and assumptions before interpreting the statistic.
Useful distinctionUnivariate: One variable: distribution, counts, centre and spread.
Visual demonstration of Correlation vs Causation
Visual demonstration: use the diagram to trace the main objects and state changes involved in Correlation vs Causation.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Pair observations correctly and handle missing…

Pair observations correctly and handle missing pairs deliberately. This is an input-preparation stage for Correlation vs Causation. Verify the relevant type, shape, units, keys, missingness or assumptions before later steps depend on them.

Input focus: confirm the data/object, units, type, shape and assumptions before the next operation depends on them.
Mathematical / formal view
Pearson r = covariance(X, Y) / (standard deviation(X) × standard deviation(Y)); therefore r is unitless and lies between -1 and 1.
How it works

Trace the mechanism step by step

  1. Pair observations correctly and handle missing pairs deliberately.
  2. Plot the variables to check shape, clusters and outliers.
  3. Compute Pearson r for linear association or Spearman rho for monotonic ranked association when appropriate.
  4. Interpret sign as direction and magnitude as strength relative to context.
  5. Check whether subgroups or influential points change the coefficient.
  6. Never turn correlation into a causal statement without a design that identifies causality.
Worked demonstration

Make the concept concrete

Demonstration

Python / pandas example

# Step 1 — Import the module so its functions/classes are available to the rest of this example.
import pandas as pd
# Step 2 — Construct `df` as a tabular object with named columns for inspectable analysis.
df = pd.DataFrame({"x":[1,2,3,4,5], "y":[2,3,5,8,12]})
# Step 3 — Display the current value explicitly so the result/state can be inspected during execution.
print("Pearson:", round(df["x"].corr(df["y"], method="pearson"), 3))
# Step 4 — Display the current value explicitly so the result/state can be inspected during execution.
print("Spearman:", round(df["x"].corr(df["y"], method="spearman"), 3))
Expected / illustrative result
Pearson: about 0.98
Spearman: 1.0
The relationship is perfectly monotonic but not perfectly linear.
Interpret the result.

For Correlation vs Causation, connect the displayed result to the specific input and mechanism above; independently verify one value/state change rather than treating successful execution as proof.

Distinctions & related ideas

Know what this is — and what it is not

UnivariateOne variable: distribution, counts, centre and spread.
BivariateTwo variables: association or group differences.
MultivariateSeveral variables: interactions, confounding, conditional patterns.
Confirmatory analysisTests or models a pre-specified claim; should be distinguished from open-ended exploration.
Use deliberately

When it is appropriate

Use Correlation vs Causation when the statistical quantity or inferential idea matches the variable type, sampling process and question being asked.

Boundary conditions

When to stop or reconsider

Do not interpret the result beyond the assumptions and design that support it; distinguish descriptive evidence, uncertainty and causal claims.

Common mistakes

Failure modes to recognise

  • Using a summary or test that does not match the variable scale, dependence structure or sampling design.
  • Treating a point estimate or p-value as a complete statement without effect size, uncertainty or context.
  • Confusing association with causation or sample behaviour with a guaranteed population truth.
Verification

How to check the result

  • Compute a tiny example by hand or simulate a simple case where the expected behaviour is known.
  • Check units, sample size, denominator and assumptions before interpreting the statistic.
  • Change one observation/assumption and predict how the estimate or uncertainty should respond.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Correlation vs Causation. First pair observations correctly and handle missing pairs deliberately. Then plot the variables to check shape, clusters and outliers. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Use a very small numeric/categorical example and calculate one quantity manually. Separate what is observed in the sample from what is inferred about a wider process.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Correlation vs Causation, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Pair observations correctly and handle missing pairs deliberately.
Step 2Plot the variables to check shape, clusters and outliers.
Step 3Compute Pearson r for linear association or Spearman rho for monotonic ranked association when appropriate.
Step 4Interpret sign as direction and magnitude as strength relative to context.
Lesson summary

What to remember

  • Correlation summarises the direction and strength of association between variables. Pearson correlation measures linear association on the original numeric scale; Spearman correlation measures monotonic association after converting values to ranks. A coefficient is not a causal estimate and should be interpreted together with the scatter plot and data-generating context.
  • Pair observations correctly and handle missing pairs deliberately.
  • Using a summary or test that does not match the variable scale, dependence structure or sampling design.
  • Compute a tiny example by hand or simulate a simple case where the expected behaviour is known.