4 · Data Cleaning & Missing Data

Duplicates, Entities & Consistency

Duplicates, Entities & Consistency groups the core ideas a learner needs at the 4 · data cleaning & missing data stage. Work through the lessons in order when new to the area, or use them independently as a reference when implementing an analysis.

How to use this topic

Learn the mechanism one decision at a time

Work through the lessons in order if the topic is new. If you already know the basics, open the specific leaf lesson that matches the operation, diagnostic or failure mode you need.

1Definition→
2Mechanism→
3Example→
4Diagnostic→
5Decision
01
Exact duplicate rowsExact duplicate rows is a practical concept within Duplicates, Entities & Consistency. It helps turn the broader workflow stage “4 · Data Cleaning & Missing Data” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.
02
Duplicate identifiersDuplicate identifiers is a practical concept within Duplicates, Entities & Consistency. It helps turn the broader workflow stage “4 · Data Cleaning & Missing Data” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.
03
Fuzzy duplicate recordsFuzzy duplicate records is a practical concept within Duplicates, Entities & Consistency. It helps turn the broader workflow stage “4 · Data Cleaning & Missing Data” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.
04
Entity resolutionEntity resolution is a practical concept within Duplicates, Entities & Consistency. It helps turn the broader workflow stage “4 · Data Cleaning & Missing Data” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.
05
Canonical categories and spellingCanonical categories and spelling is a practical concept within Duplicates, Entities & Consistency. It helps turn the broader workflow stage “4 · Data Cleaning & Missing Data” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.
06
Unit and currency consistencyUnit and currency consistency is a practical concept within Duplicates, Entities & Consistency. It helps turn the broader workflow stage “4 · Data Cleaning & Missing Data” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.
07
Cross-field consistency rulesCross-field consistency rules is a practical concept within Duplicates, Entities & Consistency. It helps turn the broader workflow stage “4 · Data Cleaning & Missing Data” into an explicit analytical decision that can be explained, implemented and checked. The concept should be understood in terms of purpose, mechanism, assumptions, evidence and downstream consequences.