- Describe how machine learning differs from explicit rules.
- Identify examples, features, and labels in a table.
- Separate fitting from prediction.
- Explain why a correlation is not a cause.
- Propose a baseline and a fair initial test.
Before you begin
Know what inputs, outputs, and models are. No programming is needed for this lesson.
A problem whose rules are hard to list
A school delivery team wants to estimate how long a cart takes to reach another building. Distance matters, but so do load, route, and interruptions. You could write rules for every situation. Alternatively, collect past trips and fit a model relating available measurements to travel time. This second approach is machine learning.
ML still uses ordinary programming. Someone must read records, choose the model, fit it, evaluate results, and display predictions. The change is that some numerical relationships are fitted from examples rather than written directly by a person. Google's introduction describes this use of data to train predictive or generative software. What is machine learning?.
ML is not necessary for every task. Calculating a bill from known prices is better handled with an exact formula. Consider learning when useful patterns exist in examples and the task is difficult to specify completely with reliable rules.
Read the training table
One row represents one completed trip. Features are the input measurements available when making a prediction: distance in meters and load in kilograms, for example. The label is the target value observed afterward: travel time in seconds. Training data contains the examples used to fit the model.
A feature must be available at prediction time. Using the actual arrival time to predict trip duration would expose information from the answer. This is leakage and can create unrealistically good evaluation results. Choose inputs based on the real sequence of events.
Features also need consistent definitions. Does distance mean straight-line distance or the actual route? Does travel time include waiting at doors? A carefully defined small dataset is more informative than a larger table whose columns change meaning.
| Trip | Route distance (m) | Load (kg) | Travel time (s) |
|---|---|---|---|
| A | 10 | 1 | 18 |
| B | 20 | 1 | 31 |
| C | 20 | 4 | 38 |
A model represents a relationship
A simple model might estimate time as a weighted combination of distance and load, plus a constant. Training chooses the weights from the recorded examples. Prediction applies the fitted relationship to a new input. The model does not need to store a separate manually written rule for every possible distance.
The resulting number is an estimate, not a guarantee. If a door closes unexpectedly, the model may have no information that predicts the delay. An error can arise from incomplete inputs as well as from a poor fitting method.
Do not infer cause from prediction alone. Heavy loads might coincide with busy school periods, so a relationship between load and delay could partly reflect congestion. A model that predicts well has not automatically established what would happen if you changed one factor experimentally.
Work through a baseline prediction
Before building a complicated model, choose a baseline: a simple comparison method. Predicting the average training trip duration is one option. It ignores distance and load, but it gives your learned model a result to beat.
For travel times 18, 31, and 38 seconds, the mean is 29 seconds. On a new trip that takes 35 seconds, the baseline's absolute error is six seconds. If your learned model predicts 33, its absolute error is two seconds on that trip. One comparison is not enough; repeat across suitable unseen trips.
Keep test examples separate while fitting. If you continually change a model after inspecting the same test cases, they become part of development. Later lessons separate training, validation, and testing more carefully.
Decide whether the result is useful
A two-second error may be acceptable for a rough classroom display but unacceptable for coordinating two moving carts. Usefulness depends on the purpose. State the relevant units, acceptable error, and conditions where the predictor should decline or warn.
Small experiments can run on an ordinary computer. Expensive hardware is not required to understand features, labels, fitting, and evaluation. You will learn more by inspecting a small dataset than by running a large model whose inputs and failures you cannot explain.
Important terms
- Example
- One observation or row used in learning or evaluation.
- Feature
- An input available for prediction.
- Label
- The target output recorded for an example.
- Training data
- Examples used to fit model parameters.
- Prediction
- A model's estimate for a supplied input.
- Baseline
- A simple method used as a comparison.
Mini project: Design a delivery dataset
- Draw a table with distance, load, and duration, including units.
- Write six invented trips and label them simulated data.
- Identify when each value becomes available.
- Reserve two trips for a later check and calculate the mean duration of the other four.
- Finish by predicting the reserved trips with the mean and calculating their absolute errors.
Common mistakes and debugging
- Using an input that reveals the answer: check what is known before prediction.
- Treating correlation as cause: consider other explanations or controlled experiments.
- Reporting training fit as proof: evaluate new examples.
- Forgetting units and definitions: document every column before fitting.
Independent challenge
Add a route-type feature. Explain how you would record it consistently and why it might help beyond distance alone.
Check your understanding: 10 questions
What changes in ML compared with explicit rules?
What is the label in the delivery example?
Why exclude actual arrival time as a feature?
What is the baseline mean of 18, 31, and 38?
Does good prediction establish causation?
In your own words, what does “Example” mean?
In your own words, what does “Feature” mean?
In your own words, what does “Label” mean?
In your own words, what does “Training data” mean?
In your own words, what does “Prediction” mean?
Quiz answers
Reveal all 10 answers after your attempt
- Some model relationships are fitted from examples rather than specified directly.
- Observed travel time.
- It is unavailable before the trip and can reveal the target.
- 29 seconds.
- No. Predictive relationships can reflect confounding factors.
- One observation or row used in learning or evaluation.
- An input available for prediction.
- The target output recorded for an example.
- Examples used to fit model parameters.
- A model's estimate for a supplied input.
Summary
ML fits models from examples. Features, labels, data definitions, a baseline, and independent tests make a prediction project understandable and honest.
Continue learning
ML02 distinguishes the main learning setups and the kinds of problems each addresses.
- How AI Models Are Trained
- Supervised, Unsupervised, and Reinforcement Learning
- Training Data: Quality, Bias, Cleaning, and Data Splits
- What Is Programming, and Why Learn Python?
Sources and further reading
Prepared 2026-09-18. Draft — terminology source checked; baseline arithmetic checked