Reporting & Reproducibility · Lesson 86

Reproducible Scripts

A reproducible analysis script can be rerun from declared inputs to regenerate the same transformations and outputs.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

Reproducible Scripts

A reproducible analysis script can be rerun from declared inputs to regenerate the same transformations and outputs. It avoids hidden notebook state and manual spreadsheet edits by making paths, parameters, dependencies and execution order explicit.

Learning goal: explain why Reproducible Scripts behaves this way, apply it to a small example, and verify the result independently. Begin by being able to justify this first step: Read inputs from declared locations rather than manually edited in-memory data.

Deeper walkthrough

Read Reproducible Scripts as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Read inputs from declared locations rather than manually edited in-memory data. Stage 2: Place transformations in deterministic functions where practical. Stage 3: Parameterise dates/paths instead of changing source code between runs. Final checkpoint: Record dependencies and rerun in a clean environment.

Mechanism

Follow the transformation

Read inputs from declared locations rather than manually edited in-memory data.

Place transformations in deterministic functions where practical.

Parameterise dates/paths instead of changing source code between runs.

Evidence

Know what would convince you

  • Recompute one result from a handful of source rows or an independent formula.
  • Check row counts, group totals and units before interpreting differences.
Useful distinctionDefinition: The exact metric/selection/comparison being computed.
Visual demonstration of Reproducible Scripts
Visual demonstration: use the diagram to trace the main objects and state changes involved in Reproducible Scripts.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Read inputs from declared locations rather…

Read inputs from declared locations rather than manually edited in-memory data. This is an input-preparation stage for Reproducible Scripts. Verify the relevant type, shape, units, keys, missingness or assumptions before later steps depend on them.

Input focus: confirm the data/object, units, type, shape and assumptions before the next operation depends on them.
How it works

Trace the mechanism step by step

  1. Read inputs from declared locations rather than manually edited in-memory data.
  2. Place transformations in deterministic functions where practical.
  3. Parameterise dates/paths instead of changing source code between runs.
  4. Write outputs to known locations with versionable names.
  5. Record dependencies and rerun in a clean environment.
Worked demonstration

Reproducible entry point

# Step 1 — Define the reusable `main` function; its indented body describes what happens for each call.
def main(input_path, output_path):
    # load -> validate -> analyse -> save
    # Step 2 — Execute this statement and inspect how it changes the current value, object or program state.
    ...

# Step 3 — Evaluate this condition and execute the indented branch only when the condition is true.
if __name__ == "__main__":
    # Step 4 — Leave this block intentionally empty as a placeholder for later behaviour.
    pass
Expected / illustrative result
A single declared entry point makes the execution order explicit; real code would fill in each validated stage.
Interpret the result.

For Reproducible Scripts, identify exactly what each reported quantity represents, including its units/denominator, and independently recompute one part of the result.

Distinctions & related ideas

Place the concept correctly

DefinitionThe exact metric/selection/comparison being computed.
EvidenceTable, formula or visual that answers the question.
AuditIndependent count/total/rule check that can reveal an error.
Use deliberately

When it is appropriate

Use Reproducible Scripts when it answers a defined question in Reporting & Reproducibility and its inputs/assumptions match the current data or program state.

Boundary conditions

When to stop or reconsider

Reconsider Reproducible Scripts when the required information is unavailable, the operation would violate a validation/data boundary, or a simpler operation answers the question more transparently.

Common mistakes

Failure modes to recognise

  • Changing the population/grain without noticing it.
  • Using an undefined denominator, time window, unit or category rule.
  • Presenting a number/plot without reconciling it to source counts or totals.
Verification

How to check the result

  • Recompute one result from a handful of source rows or an independent formula.
  • Check row counts, group totals and units before interpreting differences.
  • Change one source value and predict which reported value/mark should change.
Hands-on practice

Demonstrate understanding

Try this:

Construct a tiny example of Reproducible Scripts. First read inputs from declared locations rather than manually edited in-memory data. Then place transformations in deterministic functions where practical. Predict the result before execution and explain one boundary or failure case.

Use a very small example and calculate one quantity manually. Separate sample evidence from population/causal claims.
Knowledge check

Check reasoning, not memorisation

Which approach best demonstrates understanding of Reproducible Scripts?

Quick reference

Remember the logic

Step 1Read inputs from declared locations rather than manually edited in-memory data.
Step 2Place transformations in deterministic functions where practical.
Step 3Parameterise dates/paths instead of changing source code between runs.
Step 4Write outputs to known locations with versionable names.
Lesson summary

What to remember

  • A reproducible analysis script can be rerun from declared inputs to regenerate the same transformations and outputs. It avoids hidden notebook state and manual spreadsheet edits by making paths, parameters, dependencies and execution order explicit.
  • Read inputs from declared locations rather than manually edited in-memory data.
  • Changing the population/grain without noticing it.
  • Recompute one result from a handful of source rows or an independent formula.