Python, NumPy & pandas · Lesson 12

Dates and Categories

Dates and categorical variables need semantic dtypes because treating them as generic strings loses useful structure.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

Dates and Categories

Dates and categorical variables need semantic dtypes because treating them as generic strings loses useful structure. Parsed dates support ordering, durations and calendar features; categorical dtypes represent a finite set of labels and can distinguish unordered categories from ordered levels.

Learning goal: explain why Dates and Categories behaves this way, apply it to a small example, and verify the result independently. Begin by being able to justify this first step: Parse date strings with an explicit format or validate inferred parsing.

Deeper walkthrough

Read Dates and Categories as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Parse date strings with an explicit format or validate inferred parsing. Stage 2: Check time zone and missing/invalid dates. Stage 3: Convert low-cardinality label fields to category only when the category set is meaningful. Final checkpoint: Verify that arithmetic/comparisons match the semantic type.

Mechanism

Follow the transformation

Parse date strings with an explicit format or validate inferred parsing.

Check time zone and missing/invalid dates.

Convert low-cardinality label fields to category only when the category set is meaningful.

Evidence

Know what would convince you

  • Confirm fitted transformations/models saw only training data.
  • Retain fold/test predictions so metrics can be recomputed independently.
Useful distinctionTraining evidence: Information allowed to influence fitted state.
Click a stage to inspect what happens, what changes, and what should be checked before moving on.
Stage 1

Parse date strings with an explicit…

Parse date strings with an explicit format or validate inferred parsing. For Dates and Categories, make this checkpoint explicit by recording the evidence inspected, the expected result, and the condition that would make you reject the current result.

Verification focus: record the evidence you inspected and the condition that would make this stage fail.
How it works

Trace the mechanism step by step

  1. Parse date strings with an explicit format or validate inferred parsing.
  2. Check time zone and missing/invalid dates.
  3. Convert low-cardinality label fields to category only when the category set is meaningful.
  4. Declare ordered categories when rank is intrinsic.
  5. Verify that arithmetic/comparisons match the semantic type.
Worked demonstration

Typed columns

# Step 1 — Import the module so its functions/classes are available to the rest of this example.
import pandas as pd
# Step 2 — Construct `df` as a tabular object with named columns for inspectable analysis.
df = pd.DataFrame({"date":["2026-01-01","2026-01-02"],"level":["low","high"]})
# Step 3 — Execute this statement and inspect how it changes the current value, object or program state.
df["date"] = pd.to_datetime(df["date"])
# Step 4 — Execute this statement and inspect how it changes the current value, object or program state.
df["level"] = pd.Categorical(df["level"], categories=["low","medium","high"], ordered=True)
# Step 5 — Display the current value explicitly so the result/state can be inspected during execution.
print(df.dtypes)
Expected / illustrative result
The date becomes a datetime type and level becomes an ordered categorical variable.
Interpret the result.

For Dates and Categories, trace representative input values into the result and verify shape, dtype, row grain, axis or key behaviour that the operation can change.

Distinctions & related ideas

Place the concept correctly

Training evidenceInformation allowed to influence fitted state.
Held-out evidenceIndependent observations used to estimate generalisation.
InterpretationWhat the result supports, with assumptions and limitations.
Use deliberately

When it is appropriate

Use Dates and Categories when it answers a defined question in Python, NumPy & pandas and its inputs/assumptions match the current data or program state.

Boundary conditions

When to stop or reconsider

Reconsider Dates and Categories when the required information is unavailable, the operation would violate a validation/data boundary, or a simpler operation answers the question more transparently.

Common mistakes

Failure modes to recognise

  • Learning preprocessing/feature/model choices from held-out test information.
  • Comparing models under different splits or preprocessing and attributing the difference to the algorithm.
  • Turning an association or model explanation into an unsupported causal claim.
Verification

How to check the result

  • Confirm fitted transformations/models saw only training data.
  • Retain fold/test predictions so metrics can be recomputed independently.
  • Inspect errors/subgroups and compare with a baseline before generalising the conclusion.
Hands-on practice

Demonstrate understanding

Try this:

Construct a tiny example of Dates and Categories. First parse date strings with an explicit format or validate inferred parsing. Then check time zone and missing/invalid dates. Predict the result before execution and explain one boundary or failure case.

Use 4–8 rows containing the exact key/category/missing-value pattern. Trace one row or group from input to output.
Knowledge check

Check reasoning, not memorisation

Which approach best demonstrates understanding of Dates and Categories?

Quick reference

Remember the logic

Step 1Parse date strings with an explicit format or validate inferred parsing.
Step 2Check time zone and missing/invalid dates.
Step 3Convert low-cardinality label fields to category only when the category set is meaningful.
Step 4Declare ordered categories when rank is intrinsic.
Lesson summary

What to remember

  • Dates and categorical variables need semantic dtypes because treating them as generic strings loses useful structure. Parsed dates support ordering, durations and calendar features; categorical dtypes represent a finite set of labels and can distinguish unordered categories from ordered levels.
  • Parse date strings with an explicit format or validate inferred parsing.
  • Learning preprocessing/feature/model choices from held-out test information.
  • Confirm fitted transformations/models saw only training data.