Data Science Foundations · Lesson 5

Prediction vs Explanation vs Causality

Prediction vs Explanation vs Causality establishes the purpose and structure of a data-science project.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Prediction vs Explanation vs Causality actually means

Prediction vs Explanation vs Causality establishes the purpose and structure of a data-science project. Data science combines problem framing, data engineering, statistics, computing and modelling to produce reproducible evidence or predictions that support a real decision.

Prediction vs Explanation vs Causality matters because data science combines problem framing, data, statistics, computation and modelling. Clear purpose and reproducible structure keep the technical work connected to the real prediction, explanation or decision objective.

Deeper walkthrough

Read Prediction vs Explanation vs Causality as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Translate the real question into an analytical target or estimand. Stage 2: Define the unit of analysis and time of prediction/decision. Stage 3: Specify what information is legitimately available at that time. Final checkpoint: Evaluate using metrics that reflect the decision costs and communicate uncertainty.

Mechanism

Follow the transformation

Translate the real question into an analytical target or estimand.

Define the unit of analysis and time of prediction/decision.

Specify what information is legitimately available at that time.

Evidence

Know what would convince you

  • Write the decision question, unit of analysis and metric formula in plain language before computing it.
  • Reconcile KPI numerators/denominators or grouped totals to source counts for a small slice.
Useful distinctionExplanation: Understand relationships/mechanisms; interpretability and causal design may dominate.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Translate the real question into an…

Translate the real question into an analytical target or estimand. For Prediction vs Explanation vs Causality, identify the exact state before this stage, the operation or rule applied here, and the observable state afterwards so the mechanism remains inspectable.

State focus: identify exactly what changed at this stage and what observable evidence confirms that change.
How it works

Trace the mechanism step by step

  1. Translate the real question into an analytical target or estimand.
  2. Define the unit of analysis and time of prediction/decision.
  3. Specify what information is legitimately available at that time.
  4. Build a reproducible pipeline from raw inputs to results.
  5. Evaluate using metrics that reflect the decision costs and communicate uncertainty.
Worked demonstration

Make the concept concrete

Demonstration

Text example

Observation: umbrella sales and traffic accidents both rise on rainy days.
Correlation: umbrella sales ↔ accidents.
Confounder: rain influences both.
Causal claim “umbrellas cause accidents” is not supported by the correlation.
Expected / illustrative result
Prediction can exploit association; explanation and causal intervention questions require additional design/assumptions.
Interpret the result.

For Prediction vs Explanation vs Causality, connect the displayed result to the specific input and mechanism above; independently verify one value/state change rather than treating successful execution as proof.

Distinctions & related ideas

Know what this is — and what it is not

ExplanationUnderstand relationships/mechanisms; interpretability and causal design may dominate.
PredictionAccurately estimate unknown outcomes for new cases.
Causal inferenceEstimate effects of interventions under identification assumptions.
Data analyticsOften emphasises descriptive/diagnostic decision support; overlaps substantially with data science.
Use deliberately

When it is appropriate

Use Prediction vs Explanation vs Causality when it connects a clearly framed stakeholder question to measurable evidence and an action or decision.

Boundary conditions

When to stop or reconsider

Reframe the analysis if the decision, unit of analysis, metric definition, comparison group or time window is still ambiguous; more computation will not repair an undefined question.

Common mistakes

Failure modes to recognise

  • Starting with a favourite chart/tool before defining the decision and unit of analysis.
  • Using an undefined KPI, denominator, cohort or time window and then comparing incomparable numbers.
  • Turning association or a descriptive pattern into a causal recommendation without supporting design/evidence.
Verification

How to check the result

  • Write the decision question, unit of analysis and metric formula in plain language before computing it.
  • Reconcile KPI numerators/denominators or grouped totals to source counts for a small slice.
  • Test whether the conclusion changes under one reasonable alternative definition or comparison window.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Prediction vs Explanation vs Causality. First translate the real question into an analytical target or estimand. Then define the unit of analysis and time of prediction/decision. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Write the decision, unit and metric definition first. Use a tiny slice where you can recompute the KPI/table manually and explain what would change the recommendation.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Prediction vs Explanation vs Causality, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Translate the real question into an analytical target or estimand.
Step 2Define the unit of analysis and time of prediction/decision.
Step 3Specify what information is legitimately available at that time.
Step 4Build a reproducible pipeline from raw inputs to results.
Lesson summary

What to remember

  • Prediction vs Explanation vs Causality establishes the purpose and structure of a data-science project. Data science combines problem framing, data engineering, statistics, computing and modelling to produce reproducible evidence or predictions that support a real decision.
  • Translate the real question into an analytical target or estimand.
  • Starting with a favourite chart/tool before defining the decision and unit of analysis.
  • Write the decision question, unit of analysis and metric formula in plain language before computing it.