What you will learn
  • Identify inconsistent units and missing values.
  • Distinguish data errors from unusual valid observations.
  • Recognize sampling and labeling bias.
  • Choose random, grouped, or chronological splits.
  • Prevent preprocessing leakage.

Before you begin

Know features, labels, training, and loss. ML04 provides a companion example but is not required to audit a table.

Good data matches a defined task

A temperature table may look tidy while hiding several problems: Celsius and Fahrenheit mixed in one column, missing values stored as zero, or many readings copied from the same minute. Before choosing a model, ask what one row means and how each value was produced.

Good data is not merely large. It has appropriate coverage, consistent definitions, usable labels, and provenance: a record of where it came from. For a room-temperature predictor, data from one sunny afternoon cannot represent all seasons or rooms.

Write a data dictionary naming each column, unit, acceptable range, missing-value convention, and measurement method. Preserve the original file before cleaning so transformations are traceable and mistakes can be corrected.

Clean with reasons, not reflexes

Suppose a room sensor reports 21.4, 21.5, blank, and 87.0 degrees Celsius. Blank means missing, not freezing. The high reading might be a device fault, a unit mix-up, or a sensor placed near a heat source. Inspect context before deleting it.

Imputation replaces missing values using a rule, such as a training-set median. It can help algorithms that do not accept missing entries, but it does not recover the actual measurement. Add a missingness indicator when absence itself may matter, and document the assumption.

Removing unusual observations automatically can erase the rare cases you most need to detect. Separate invalid records from valid extremes. A value can be uncommon yet physically meaningful; another can be ordinary-looking but assigned to the wrong timestamp.

Bias enters through collection and labels

If a voice dataset contains mostly quiet recordings from one age group, evaluation on similar recordings can hide failure elsewhere. Sampling bias occurs when collected examples fail to represent the intended population or conditions. Labeling bias can arise when annotators interpret categories inconsistently.

More data collected through the same narrow process may preserve the same gap. Ask who or what is absent, which conditions are uncommon, and whether errors differ across relevant groups. Do not infer that one global average establishes equal performance.

A target can also be a proxy: an available measurement standing in for the outcome you actually care about. Predicting past decisions does not prove those decisions were correct. Record what labels mean and what they leave unmeasured.

Split for the future you intend to predict

Training data fits parameters; validation data guides development; test data assesses the selected approach. The right split depends on how examples relate. Independent shuffled examples may suit a random split, while repeated records from one person or device should often stay together.

For forecasting, chronological splits usually better reflect predicting later data from earlier information. Randomly scattering future observations into training can make the problem unrealistically easy. A group split can reserve entire devices or trips to assess transfer to new ones.

Keep duplicates and near-duplicates from crossing boundaries. Otherwise, recognizing a nearly repeated example may masquerade as generalization. The test should resemble the actual challenge, not merely satisfy a percentage recipe.

Fit preprocessing on training data only

If you calculate a mean or median using the entire dataset before splitting, test information influences the pipeline. Even without fitting the final model on test labels, you have allowed leakage through preparation.

Split first. Learn imputation, scaling, and other fitted transformations from training data, then apply those same transformations to validation and test data. A pipeline helps keep these steps together. scikit-learn's common pitfalls guide explains this prevention pattern.

For our sensor table, record missing counts and units first, reserve later days for evaluation, fit any replacement rule on earlier training days, and retain a log of excluded records. The finished deliverable is a defensible dataset process, not a spreadsheet with every suspicious cell hidden.

Important terms

Provenance
The source and history of data.
Imputation
Replacing missing values using a documented rule.
Sampling bias
A collection process that poorly represents intended use.
Proxy label
An available target standing in for another desired outcome.
Leakage
Information entering development that would not be available legitimately.
Pipeline
A linked sequence of preparation and modeling steps.

Mini project: Audit ten sensor rows

  1. Create ten simulated rows with timestamp, device, temperature, and unit.
  2. Include a blank, duplicate, unit mismatch, and unusual value.
  3. Write an action and reason for each issue; preserve originals.
  4. Propose a chronological split and identify which cleaning decisions use learned statistics.
  5. Finish with a data dictionary and audit log someone else can follow.

Common mistakes and debugging

  • Replacing every blank with zero: distinguish missing from measured zero.
  • Deleting every outlier: investigate valid extremes.
  • Randomly splitting related records: group or order them to match intended use.
  • Fitting preparation on all data: learn transformations from training data only.

Independent challenge

Suppose one device records only in warm rooms. Explain why device ID might become a shortcut and design a test that reveals it.

Check your understanding: 10 questions

  1. Does blank temperature mean zero?

  2. What does imputation recover?

  3. Why group records by device when testing new devices?

  4. Where should a scaling mean be fitted?

  5. Can more data preserve bias?

  6. In your own words, what does “Provenance” mean?

  7. In your own words, what does “Imputation” mean?

  8. In your own words, what does “Sampling bias” mean?

  9. In your own words, what does “Proxy label” mean?

  10. In your own words, what does “Leakage” mean?

Quiz answers

Reveal all 10 answers after your attempt
  1. No. Missingness and a measured zero are different.
  2. A replacement based on an assumption, not the true missing measurement.
  3. To avoid having the same device's patterns on both sides of evaluation.
  4. On training data only, then reused for other sets.
  5. Yes, if the collection process repeats the same missing coverage or labeling problems.
  6. The source and history of data.
  7. Replacing missing values using a documented rule.
  8. A collection process that poorly represents intended use.
  9. An available target standing in for another desired outcome.
  10. Information entering development that would not be available legitimately.

Summary

Define, preserve, inspect, and split data before fitting. Honest preparation includes missingness, coverage, and leakage checks rather than cosmetic cleaning alone.

Continue learning

ML06 explains neural networks while keeping these data and evaluation principles in place.

Choose a connected learning path

Sources and further reading

Prepared 2026-09-18. Draft — data-handling concepts; example audit is simulated