- Identify inconsistent units and missing values.
- Distinguish data errors from unusual valid observations.
- Recognize sampling and labeling bias.
- Choose random, grouped, or chronological splits.
- Prevent preprocessing leakage.
Before you begin
Know features, labels, training, and loss. ML04 provides a companion example but is not required to audit a table.
Good data matches a defined task
A temperature table may look tidy while hiding several problems: Celsius and Fahrenheit mixed in one column, missing values stored as zero, or many readings copied from the same minute. Before choosing a model, ask what one row means and how each value was produced.
Good data is not merely large. It has appropriate coverage, consistent definitions, usable labels, and provenance: a record of where it came from. For a room-temperature predictor, data from one sunny afternoon cannot represent all seasons or rooms.
Write a data dictionary naming each column, unit, acceptable range, missing-value convention, and measurement method. Preserve the original file before cleaning so transformations are traceable and mistakes can be corrected.
Clean with reasons, not reflexes
Suppose a room sensor reports 21.4, 21.5, blank, and 87.0 degrees Celsius. Blank means missing, not freezing. The high reading might be a device fault, a unit mix-up, or a sensor placed near a heat source. Inspect context before deleting it.
Imputation replaces missing values using a rule, such as a training-set median. It can help algorithms that do not accept missing entries, but it does not recover the actual measurement. Add a missingness indicator when absence itself may matter, and document the assumption.
Removing unusual observations automatically can erase the rare cases you most need to detect. Separate invalid records from valid extremes. A value can be uncommon yet physically meaningful; another can be ordinary-looking but assigned to the wrong timestamp.
Bias enters through collection and labels
If a voice dataset contains mostly quiet recordings from one age group, evaluation on similar recordings can hide failure elsewhere. Sampling bias occurs when collected examples fail to represent the intended population or conditions. Labeling bias can arise when annotators interpret categories inconsistently.
More data collected through the same narrow process may preserve the same gap. Ask who or what is absent, which conditions are uncommon, and whether errors differ across relevant groups. Do not infer that one global average establishes equal performance.
A target can also be a proxy: an available measurement standing in for the outcome you actually care about. Predicting past decisions does not prove those decisions were correct. Record what labels mean and what they leave unmeasured.
Split for the future you intend to predict
Training data fits parameters; validation data guides development; test data assesses the selected approach. The right split depends on how examples relate. Independent shuffled examples may suit a random split, while repeated records from one person or device should often stay together.
For forecasting, chronological splits usually better reflect predicting later data from earlier information. Randomly scattering future observations into training can make the problem unrealistically easy. A group split can reserve entire devices or trips to assess transfer to new ones.
Keep duplicates and near-duplicates from crossing boundaries. Otherwise, recognizing a nearly repeated example may masquerade as generalization. The test should resemble the actual challenge, not merely satisfy a percentage recipe.
Fit preprocessing on training data only
If you calculate a mean or median using the entire dataset before splitting, test information influences the pipeline. Even without fitting the final model on test labels, you have allowed leakage through preparation.
Split first. Learn imputation, scaling, and other fitted transformations from training data, then apply those same transformations to validation and test data. A pipeline helps keep these steps together. scikit-learn's common pitfalls guide explains this prevention pattern.
For our sensor table, record missing counts and units first, reserve later days for evaluation, fit any replacement rule on earlier training days, and retain a log of excluded records. The finished deliverable is a defensible dataset process, not a spreadsheet with every suspicious cell hidden.
Important terms
- Provenance
- The source and history of data.
- Imputation
- Replacing missing values using a documented rule.
- Sampling bias
- A collection process that poorly represents intended use.
- Proxy label
- An available target standing in for another desired outcome.
- Leakage
- Information entering development that would not be available legitimately.
- Pipeline
- A linked sequence of preparation and modeling steps.
Mini project: Audit ten sensor rows
- Create ten simulated rows with timestamp, device, temperature, and unit.
- Include a blank, duplicate, unit mismatch, and unusual value.
- Write an action and reason for each issue; preserve originals.
- Propose a chronological split and identify which cleaning decisions use learned statistics.
- Finish with a data dictionary and audit log someone else can follow.
Common mistakes and debugging
- Replacing every blank with zero: distinguish missing from measured zero.
- Deleting every outlier: investigate valid extremes.
- Randomly splitting related records: group or order them to match intended use.
- Fitting preparation on all data: learn transformations from training data only.
Independent challenge
Suppose one device records only in warm rooms. Explain why device ID might become a shortcut and design a test that reveals it.
Check your understanding: 10 questions
Does blank temperature mean zero?
What does imputation recover?
Why group records by device when testing new devices?
Where should a scaling mean be fitted?
Can more data preserve bias?
In your own words, what does “Provenance” mean?
In your own words, what does “Imputation” mean?
In your own words, what does “Sampling bias” mean?
In your own words, what does “Proxy label” mean?
In your own words, what does “Leakage” mean?
Quiz answers
Reveal all 10 answers after your attempt
- No. Missingness and a measured zero are different.
- A replacement based on an assumption, not the true missing measurement.
- To avoid having the same device's patterns on both sides of evaluation.
- On training data only, then reused for other sets.
- Yes, if the collection process repeats the same missing coverage or labeling problems.
- The source and history of data.
- Replacing missing values using a documented rule.
- A collection process that poorly represents intended use.
- An available target standing in for another desired outcome.
- Information entering development that would not be available legitimately.
Summary
Define, preserve, inspect, and split data before fitting. Honest preparation includes missingness, coverage, and leakage checks rather than cosmetic cleaning alone.
Continue learning
ML06 explains neural networks while keeping these data and evaluation principles in place.
- Your First Python Machine-Learning Project
- Neural Networks for Beginners
- Working With Data Using NumPy and pandas
- How to Design and Improve a Machine-Learning Project
Sources and further reading
Prepared 2026-09-18. Draft — data-handling concepts; example audit is simulated