Exploratory Data Analysis · Lesson 54

Frequency Tables

Frequency Tables is an exploratory data-analysis technique.

ConceptWorked examplePracticeKnowledge check
Textbook walkthrough

What Frequency Tables actually means

Frequency Tables is an exploratory data-analysis technique. EDA is structured investigation of distributions, groups, relationships, missingness and anomalies before stronger inferential or predictive claims are made.

Frequency Tables matters because exploratory analysis is where structure, anomalies and plausible relationships become visible before stronger claims are made. The goal is to generate and test questions while preserving uncertainty and data-quality context.

Deeper walkthrough

Read Frequency Tables as a mechanism, not a recipe

Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Start with the unit of analysis and data-quality profile. Stage 2: Use univariate summaries to inspect centre, spread, shape and unusual values. Stage 3: Compare groups using both counts and distributions rather than only averages. Final checkpoint: Record observations as hypotheses or data-quality findings, not as automatic causal claims.

Mechanism

Follow the transformation

Start with the unit of analysis and data-quality profile.

Use univariate summaries to inspect centre, spread, shape and unusual values.

Compare groups using both counts and distributions rather than only averages.

Evidence

Know what would convince you

  • Compare row/column counts, dtypes and missing values before and after the operation.
  • Trace a few representative rows or one group manually from source values to result.
Useful distinctionUnivariate: One variable: distribution, counts, centre and spread.
How it works

Trace the mechanism step by step

  1. Start with the unit of analysis and data-quality profile.
  2. Use univariate summaries to inspect centre, spread, shape and unusual values.
  3. Compare groups using both counts and distributions rather than only averages.
  4. Inspect bivariate relationships with plots and appropriate association measures.
  5. Look for missingness, clusters, nonlinear structure, time effects and subgroup patterns that could change a conclusion.
  6. Record observations as hypotheses or data-quality findings, not as automatic causal claims.
Worked demonstration

Make the concept concrete

Demonstration

Python / pandas example

# Step 1 — Import the module so its functions/classes are available to the rest of this example.
import pandas as pd
# Step 2 — Compute the right-hand expression and store its result in `s` for the next step.
s = pd.Series(["A","B","A","C","A","B"])
# Step 3 — Display the current value explicitly so the result/state can be inspected during execution.
print(s.value_counts())
# Step 4 — Display the current value explicitly so the result/state can be inspected during execution.
print((s.value_counts(normalize=True)*100).round(1))
Expected / illustrative result
Counts: A=3, B=2, C=1. Percentages: A=50.0%, B=33.3%, C=16.7%.
Interpret the result.

For Frequency Tables, trace representative source rows/columns into the result and reconcile row counts, dtypes, keys or missing values that the operation could change.

Distinctions & related ideas

Know what this is — and what it is not

UnivariateOne variable: distribution, counts, centre and spread.
BivariateTwo variables: association or group differences.
MultivariateSeveral variables: interactions, confounding, conditional patterns.
Confirmatory analysisTests or models a pre-specified claim; should be distinguished from open-ended exploration.
Use deliberately

When it is appropriate

Use Frequency Tables when the data are naturally tabular and row grain, column meaning, keys and dtypes can be stated explicitly.

Boundary conditions

When to stop or reconsider

Reconsider the operation if row identity/grain is unclear, join keys are not validated, chained transformations hide state, or the task is better expressed with a simpler table operation.

Common mistakes

Failure modes to recognise

  • Changing row grain or row count without noticing it.
  • Joining/grouping on keys whose uniqueness or missingness was never checked.
  • Interpreting a derived column or aggregation without reconciling it to source rows and units.
Verification

How to check the result

  • Compare row/column counts, dtypes and missing values before and after the operation.
  • Trace a few representative rows or one group manually from source values to result.
  • For joins/reshapes/grouping, verify key uniqueness/cardinality and reconcile totals where totals should be preserved.
Hands-on practice

Demonstrate understanding

Try this:

Build a tiny, inspectable example of Frequency Tables. First start with the unit of analysis and data-quality profile. Then use univariate summaries to inspect centre, spread, shape and unusual values. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.

Work with 4–8 rows that contain the exact key/category/missing-value pattern you want to understand. Trace one row or group all the way through.
Knowledge check

Check reasoning, not memorisation

Before trusting a result from Frequency Tables, which check provides the strongest evidence that you understand and applied it correctly?

Quick reference

Keep the important distinctions visible

Step 1Start with the unit of analysis and data-quality profile.
Step 2Use univariate summaries to inspect centre, spread, shape and unusual values.
Step 3Compare groups using both counts and distributions rather than only averages.
Step 4Inspect bivariate relationships with plots and appropriate association measures.
Lesson summary

What to remember

  • Frequency Tables is an exploratory data-analysis technique. EDA is structured investigation of distributions, groups, relationships, missingness and anomalies before stronger inferential or predictive claims are made.
  • Start with the unit of analysis and data-quality profile.
  • Changing row grain or row count without noticing it.
  • Compare row/column counts, dtypes and missing values before and after the operation.