Reading & Understanding Data · Lesson 29

Schema and Data Dictionary

A schema specifies structural expectations such as field names, types, nullability and keys.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

Schema and Data Dictionary

A schema specifies structural expectations such as field names, types, nullability and keys. A data dictionary adds semantic meaning: definition, units, allowed values, collection rules and business interpretation. Together they prevent a column name such as “rate” or “date” from being used without knowing what it actually measures.

Learning goal: explain why Schema and Data Dictionary behaves this way, apply it to a small example, and verify the result independently. Begin by being able to justify this first step: Record one-row grain and primary/unique identifiers.

Deeper walkthrough

Read Schema and Data Dictionary as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Record one-row grain and primary/unique identifiers. Stage 2: Specify each field name, logical type and nullability. Stage 3: Define units, categories and valid ranges. Final checkpoint: Validate incoming data against the schema and update the dictionary when definitions change.

Mechanism

Follow the transformation

Record one-row grain and primary/unique identifiers.

Specify each field name, logical type and nullability.

Define units, categories and valid ranges.

Evidence

Know what would convince you

  • Recompute one result from a handful of source rows or an independent formula.
  • Check row counts, group totals and units before interpreting differences.
Useful distinctionDefinition: The exact metric/selection/comparison being computed.
How it works

Trace the mechanism step by step

  1. Record one-row grain and primary/unique identifiers.
  2. Specify each field name, logical type and nullability.
  3. Define units, categories and valid ranges.
  4. Document how derived fields are calculated.
  5. Validate incoming data against the schema and update the dictionary when definitions change.
Worked demonstration

Field definition

Field: conversion_rate
Type: decimal
Definition: conversions / eligible_visits
Range: 0..1
Null rule: null when eligible_visits = 0
Expected / illustrative result
The dictionary makes the metric computable and interpretable, not just syntactically typed.
Interpret the result.

For Schema and Data Dictionary, trace representative input values into the result and verify shape, dtype, row grain, axis or key behaviour that the operation can change.

Distinctions & related ideas

Place the concept correctly

DefinitionThe exact metric/selection/comparison being computed.
EvidenceTable, formula or visual that answers the question.
AuditIndependent count/total/rule check that can reveal an error.
Use deliberately

When it is appropriate

Use Schema and Data Dictionary when it answers a defined question in Reading & Understanding Data and its inputs/assumptions match the current data or program state.

Boundary conditions

When to stop or reconsider

Reconsider Schema and Data Dictionary when the required information is unavailable, the operation would violate a validation/data boundary, or a simpler operation answers the question more transparently.

Common mistakes

Failure modes to recognise

  • Changing the population/grain without noticing it.
  • Using an undefined denominator, time window, unit or category rule.
  • Presenting a number/plot without reconciling it to source counts or totals.
Verification

How to check the result

  • Recompute one result from a handful of source rows or an independent formula.
  • Check row counts, group totals and units before interpreting differences.
  • Change one source value and predict which reported value/mark should change.
Hands-on practice

Demonstrate understanding

Try this:

Construct a tiny example of Schema and Data Dictionary. First record one-row grain and primary/unique identifiers. Then specify each field name, logical type and nullability. Predict the result before execution and explain one boundary or failure case.

Use 4–8 rows containing the exact key/category/missing-value pattern. Trace one row or group from input to output.
Knowledge check

Check reasoning, not memorisation

Which approach best demonstrates understanding of Schema and Data Dictionary?

Quick reference

Remember the logic

Step 1Record one-row grain and primary/unique identifiers.
Step 2Specify each field name, logical type and nullability.
Step 3Define units, categories and valid ranges.
Step 4Document how derived fields are calculated.
Lesson summary

What to remember

  • A schema specifies structural expectations such as field names, types, nullability and keys. A data dictionary adds semantic meaning: definition, units, allowed values, collection rules and business interpretation. Together they prevent a column name such as “rate” or “date” from being used without knowing what it actually measures.
  • Record one-row grain and primary/unique identifiers.
  • Changing the population/grain without noticing it.
  • Recompute one result from a handful of source rows or an independent formula.