- Calculate predictions from a linear model.
- Compute mean squared error.
- Explain an optimization update.
- Separate parameters and hyperparameters.
- Identify underfitting and overfitting from evidence.
Before you begin
Know features and labels. Multiplication, subtraction, and averaging are enough for the worked example.
Write the model before training it
Use a fictional conveyor experiment where input x is distance in meters and target y is travel time in seconds. A linear model predicts y = w × x + b. Weight w controls how strongly distance affects predicted time; bias b shifts all predictions. Fitting chooses these two parameter values.
The model family already imposes a limitation: it represents a straight-line relationship. It cannot capture every possible mechanical behavior. Training chooses values within that family; it does not automatically replace the family with an arbitrary explanation.
Our three constructed examples are (1, 3), (2, 5), and (3, 7). Their regularity makes arithmetic easy. Real measurements include noise and other influences, so a perfect fit here should not be confused with a realistic performance claim.
Measure the current error
Start with w = 1 and b = 0. Predictions are 1, 2, and 3. Subtract the actual labels to obtain errors −2, −3, and −4. Squaring them gives 4, 9, and 16. Their mean is 29/3, approximately 9.67.
This mean squared error, or MSE, is our loss. Squaring prevents positive and negative errors from canceling and gives larger errors more influence. Because the labels are seconds, squared errors are measured in square-seconds. Mean absolute error instead remains in seconds and is often easier to explain to a user.
Loss is a training objective; an evaluation metric is a measurement used to judge performance. The same formula can serve both roles, but you should still identify which data it is computed on.
| Distance | Observed | Prediction | Squared error |
|---|---|---|---|
| 1 | 3 | 1 | 4 |
| 2 | 5 | 2 | 9 |
| 3 | 7 | 3 | 16 |
Change parameters with a reason
Try w = 2 and b = 0. Predictions become 2, 4, and 6. Each error is −1, so MSE becomes 1. Changing b to 1 then gives 3, 5, and 7, with MSE 0. We selected these candidates by inspection; a real optimizer uses a systematic procedure.
Gradient descent estimates the local direction in parameter space that lowers loss and takes a step in that direction. The learning rate sets the step scale. Too-large updates can move past useful values; tiny updates can make progress slow.
Weight and bias are parameters learned during fitting. Learning rate, model architecture, and other chosen training settings are hyperparameters. Their selection should use development evidence, not repeated access to a final test set. Google's linear-regression course develops these relationships in greater detail.
Repeat with batches, then inspect curves
A training loop calculates predictions, measures loss, computes updates, and repeats. A batch groups examples for an update. An epoch is a pass through the training data. The exact update procedure differs by model and training method; not every algorithm uses gradient descent.
Plotting training and validation loss against progress helps diagnose behavior. If both remain poor, the model or inputs may be inadequate, or fitting may be unfinished. If training improves while validation deteriorates, overfitting is a concern. Curves are clues, not automatic diagnoses: data errors and distribution differences can produce confusing patterns.
Stopping at the lowest training loss can therefore be a bad choice. Development should consider performance on relevant examples not used for parameter updates.
A better fit is not a complete solution
Our final line predicts time 21 seconds for distance 10 meters. Nothing in the three examples establishes that this extrapolation is valid. A longer route could include a turn, a pause, or speed changes absent from training.
Keep the learning loop inside a wider process: define the task, inspect data, choose a baseline, fit, evaluate, and document limits. The optimizer handles one part. It does not decide whether the labels are meaningful or whether predictions are safe to use.
Important terms
- Weight
- A coefficient scaling an input's contribution.
- Bias
- An additive model offset.
- MSE
- The average squared prediction error.
- Optimizer
- A procedure for improving model parameters against an objective.
- Hyperparameter
- A setting selected for model or training design.
- Extrapolation
- Prediction outside the observed input range.
Mini project: Calculate a loss table
- Use the three pairs in the lesson.
- Calculate predictions for w = 1.5 and b = 1.
- Find signed errors, square them, and average.
- Check that predictions are 2.5, 4, 5.5 and MSE is 3.5/3, about 1.17.
- Finish by comparing this candidate with the w = 2, b = 0 candidate and explaining the result.
Common mistakes and debugging
- Averaging signed errors and missing cancellation: use an appropriate loss.
- Confusing MSE units with target units: squared errors have squared units.
- Calling hyperparameters learned weights: distinguish chosen settings from fitted parameters.
- Assuming training perfection ensures extrapolation: test relevant new conditions.
Independent challenge
Change the middle label from 5 to 6. Compute the loss of w = 2, b = 1 and explain why a straight line no longer matches all three points exactly.
Check your understanding: 10 questions
What does b change?
What is the initial MSE?
Why square errors?
Is learning rate a model weight?
What does falling training loss with rising validation loss suggest?
In your own words, what does “Weight” mean?
In your own words, what does “Bias” mean?
In your own words, what does “MSE” mean?
In your own words, what does “Optimizer” mean?
In your own words, what does “Hyperparameter” mean?
Quiz answers
Reveal all 10 answers after your attempt
- It shifts every prediction by the same amount.
- 29/3, approximately 9.67.
- To avoid sign cancellation and emphasize larger errors.
- No. It is a training hyperparameter.
- Possible overfitting, which should be investigated with the data and training setup.
- A coefficient scaling an input's contribution.
- An additive model offset.
- The average squared prediction error.
- A procedure for improving model parameters against an objective.
- A setting selected for model or training design.
Summary
Training changes parameters to reduce an objective. Correct arithmetic, appropriate validation, and awareness of model limits are as important as the update rule.
Continue learning
ML04 implements a small regression experiment in Python with a held-out test and baseline.
- How AI Models Are Trained
- Your First Python Machine-Learning Project
- Training Data: Quality, Bias, Cleaning, and Data Splits
- Neural Networks for Beginners
Sources and further reading
Prepared 2026-09-18. Draft — worked arithmetic checked; optimizer explanation source checked