Capstone: A Small Data Project · Lesson 162

Clean and Transform Records

Cleaning converts valid-but-messy source values into a consistent analytical representation while preserving lineage from raw to derived data.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

Clean and Transform Records

Cleaning converts valid-but-messy source values into a consistent analytical representation while preserving lineage from raw to derived data.

Learning goal: explain why Clean and Transform Records behaves this way, apply it to a small example, and verify the result independently. Begin by being able to justify this first step: Normalise categories/text, handle missing values deliberately, create derived fields with named rules and keep the raw source unchanged.

Deeper walkthrough

Read Clean and Transform Records as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Normalise categories/text, handle missing values deliberately, create derived fields with named rules and keep the raw source unchanged. Stage 2: Keep the stage inside the same data/validation definitions used by the rest of the project. Stage 3: Save the evidence produced by this stage so the next stage can be audited.

Mechanism

Follow the transformation

Normalise categories/text, handle missing values deliberately, create derived fields with named rules and keep the raw source unchanged.

Keep the stage inside the same data/validation definitions used by the rest of the project.

Save the evidence produced by this stage so the next stage can be audited.

Evidence

Know what would convince you

  • Trace a tiny input by hand and compare the runtime result.
  • Inspect type, value/shape and any mutation/side effect explicitly.
Useful distinctionInput: Objects/values supplied to the operation.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Normalise categories/text

Normalise categories/text, handle missing values deliberately, create derived fields with named rules and keep the raw source unchanged.

State focus: identify exactly what changed at this stage and what observable evidence confirms that change.
How it works

Trace the mechanism step by step

  1. Normalise categories/text, handle missing values deliberately, create derived fields with named rules and keep the raw source unchanged.
  2. Keep the stage inside the same data/validation definitions used by the rest of the project.
  3. Save the evidence produced by this stage so the next stage can be audited.
Worked demonstration

Clean and Transform Records evidence

Clean and Transform Records evidence
Evidence: before/after examples plus counts of changed, missing or rejected records.
Expected / illustrative result
The worked evidence makes the output of this project stage concrete and auditable.
Interpret the result.

For Clean and Transform Records, trace the specific input through the mechanism above and independently verify one returned value, state change or side effect.

Distinctions & related ideas

Place the concept correctly

InputObjects/values supplied to the operation.
StateNames or mutable objects that may change during execution.
OutputReturned value, side effect, file, plot or exception to inspect.
Use deliberately

When it is appropriate

Use Clean and Transform Records when it answers a defined question in Capstone: A Small Data Project and its inputs/assumptions match the current data or program state.

Boundary conditions

When to stop or reconsider

Reconsider Clean and Transform Records when the required information is unavailable, the operation would violate a validation/data boundary, or a simpler operation answers the question more transparently.

Common mistakes

Failure modes to recognise

  • Running the operation on the wrong object/type or in the wrong environment.
  • Inferring correctness from “no exception” without checking the produced value/state.
  • Hiding a boundary case instead of making its behaviour explicit.
Verification

How to check the result

  • Trace a tiny input by hand and compare the runtime result.
  • Inspect type, value/shape and any mutation/side effect explicitly.
  • Run an edge or invalid case and confirm the exception/behaviour is deliberate.
Hands-on practice

Demonstrate understanding

Try this:

Construct a tiny example of Clean and Transform Records. First normalise categories/text, handle missing values deliberately, create derived fields with named rules and keep the raw source unchanged. Then keep the stage inside the same data/validation definitions used by the rest of the project. Predict the result before execution and explain one boundary or failure case.

List the stage inputs and expected artifact, rerun it from a clean state, and compare against a concrete acceptance check.
Knowledge check

Check reasoning, not memorisation

Which approach best demonstrates understanding of Clean and Transform Records?

Quick reference

Remember the logic

Step 1Normalise categories/text, handle missing values deliberately, create derived fields with named rules and keep the raw source unchanged.
Step 2Keep the stage inside the same data/validation definitions used by the rest of the project.
Step 3Save the evidence produced by this stage so the next stage can be audited.
Lesson summary

What to remember

  • Cleaning converts valid-but-messy source values into a consistent analytical representation while preserving lineage from raw to derived data.
  • Normalise categories/text, handle missing values deliberately, create derived fields with named rules and keep the raw source unchanged.
  • Running the operation on the wrong object/type or in the wrong environment.
  • Trace a tiny input by hand and compare the runtime result.