What you will learn
  • Trace a training loop.
  • Explain a loss value.
  • Distinguish parameters from settings.
  • Explain why training performance is insufficient.
  • Separate training from retrieval.

Before you begin

Know training versus inference and context versus parameters.

Begin with examples and an objective

Imagine predicting how long a cart takes to travel a marked distance. Collect paired measurements: distance and travel time. A model receives distance and predicts time. Training needs a way to compare its prediction with the recorded result. Without an objective, learn from data does not tell us what improvement means.

For a linear model, predicted time could be weight × distance + bias. Weight changes the slope; bias shifts the prediction. These are parameters. Initial values may predict badly, and the fitting procedure adjusts them using examples.

Good training begins before adjustment. Define units and how measurements were produced. If some times include a long pause and others do not, the target has inconsistent meaning. Computation cannot repair an undefined task.

Calculate error and loss

If recorded time is eight seconds and prediction is six, the signed error is negative two seconds. Squaring it gives four square-seconds. Mean squared error averages squared errors across examples. It is one possible loss: a quantity fitting tries to reduce.

Other losses serve other tasks. Classification often uses losses based on predicted category probabilities. Language training commonly rewards probability assigned to actual next or missing tokens. The formula should match the task.

Lower loss does not satisfy every human goal. Average performance can improve while rare conditions worsen. Choose separate evaluation measures for the intended use, not only the training objective.

Examples → predictions → loss → parameter update
                ↑                    │
                └──── next pass ─────┘

Updates are guided, not magical

Gradient descent uses information about how loss changes with parameters. It adjusts them in a direction intended to reduce loss. A learning rate controls step size. Large steps can overshoot; tiny steps can progress slowly.

A batch is a group of examples used in an update. An epoch is a pass through training data. More epochs are not automatically better: a flexible model can keep fitting quirks while becoming less useful on new examples.

Walking downhill is a useful analogy, but actual optimization can involve many dimensions and noisy updates. It does not guarantee the best possible model. Other fitting procedures exist, so gradient descent is not the definition of all ML.

Keep an honest examination

Training data fits parameters. Validation data guides model settings or stopping points. Test data is reserved for assessment after those choices. Repeatedly choosing changes from test results turns the test into development information.

Cart measurements from one trip are closely related. Random splitting can put near-duplicates on both sides. Testing separate trips or conditions may better represent the intended use. Choose splits based on how future examples arise.

Overfitting means fitting training-specific patterns that fail to generalize usefully. Inspect appropriate held-out results rather than equating low training loss with success. Google: Overfitting.

Connect small training to large models

Pretraining develops broad behavior from a large learning task. Additional training can adapt behavior to instructions, domains, or preferences. Fine-tuning changes parameters; providing a document in a prompt ordinarily does not. Retrieval supplies information at inference time.

You need not train a large language model from scratch to learn these ideas. Small models make the process inspectable. Later, pretrained capabilities can help with images or language while you focus on evaluation and integration. Scale does not remove the need for a clear task and honest tests.

Important terms

Loss
A numerical measure used during fitting.
Gradient descent
Updating parameters using loss gradients.
Learning rate
A setting controlling update size.
Epoch
A pass through training data.
Validation set
Data guiding development choices.
Overfitting
Fitting patterns that fail to generalize usefully.

Mini project: Compare two models by hand

  1. Use distances 1, 2, 3 meters and times 3, 5, 7 seconds.
  2. Compare A: time = 2 × distance, and B: time = 2 × distance + 1.
  3. Calculate predictions, signed errors, squared errors, and mean squared error.
  4. Confirm MSE is 1 for A and 0 for B.
  5. Finish by explaining why perfect results on three constructed examples do not prove real reliability.

Common mistakes and debugging

  • Using training results as final evidence: reserve independent evaluation data.
  • Treating extra epochs as always helpful: monitor validation behavior.
  • Calling prompts fine-tuning: distinguish context from parameters.
  • Forgetting squared-error units differ from measurement units.

Independent challenge

Add four meters taking ten seconds. Recalculate both losses and explain why the previously perfect line is now imperfect.

Check your understanding: 10 questions

  1. Which values are learned in the linear model?

  2. What is squared error for prediction six and observation eight?

  3. What does learning rate control?

  4. Why reserve test data?

  5. Does retrieval ordinarily change parameters?

  6. In your own words, what does “Loss” mean?

  7. In your own words, what does “Gradient descent” mean?

  8. In your own words, what does “Learning rate” mean?

  9. In your own words, what does “Epoch” mean?

  10. In your own words, what does “Validation set” mean?

Quiz answers

Reveal all 10 answers after your attempt
  1. The weight and bias.
  2. Four, because (6 − 8) squared is 4.
  3. Parameter-update step size.
  4. To assess a model without using the assessment to choose it.
  5. No; it supplies information during inference.
  6. A numerical measure used during fitting.
  7. Updating parameters using loss gradients.
  8. A setting controlling update size.
  9. A pass through training data.
  10. Data guiding development choices.

Summary

Training adjusts parameters to reduce loss. Data quality and independent evaluation determine whether the resulting behavior is useful beyond training examples.

Continue learning

AI07 explains how trained models use tools in multi-step workflows.

Choose a connected learning path

Sources and further reading

Prepared 2026-09-18. Draft — concepts source checked; hand calculations checked