What you will learn
  • Write a measurable ML problem statement.
  • Select metrics matching the use case.
  • Group errors to choose improvements.
  • Run controlled development experiments.
  • Reserve final testing and plan monitoring.

Before you begin

Complete a small fitted model and understand leakage, splits, and overfitting.

Describe the decision the model supports

Suppose a school greenhouse dashboard predicts whether a plant pot will need attention tomorrow. Need attention is vague. Define an observable target, such as whether the next morning's measured soil moisture falls below a documented threshold. This is a teaching example, not advice about a particular plant species.

Identify who uses the result and what they do next. A suggestion reviewed by a teacher has different requirements from a system that automatically waters plants. The model's output, surrounding decision rule, and physical action should be specified separately.

Write success criteria before fitting: performance against a baseline, acceptable kinds of errors, usable delay, and behavior when inputs are missing. A system that cannot receive tomorrow's forecast should not depend silently on that feature.

Prepare data and a baseline

Collect data relevant to the intended conditions, record units, and preserve provenance. Choose the split based on future use: later dates, different pots, or different greenhouses may each test a different claim. Document which claim your evaluation supports.

A baseline could predict tomorrow resembles today or use a fixed recent average. It should be legitimate at prediction time and simple enough to understand. A learned method earns complexity by providing useful improvement, not merely by existing.

Build preparation and modeling into a consistent pipeline. Fit transformations on training data and keep the same feature order at inference. These details are often more important than trying another model family immediately.

Choose measurements that expose mistakes

For numerical prediction, inspect MAE and large individual errors. For a category task, count false positives and false negatives. Precision asks how many predicted positives were actually positive; recall asks how many actual positives were found. Both require a clear definition of positive.

Suppose attention is needed on only five of one hundred days. Predicting no attention every day gives 95 percent accuracy while missing every positive case. Accuracy alone hides the failure. Compare metrics with the actual consequences and frequency of outcomes.

Use a confusion matrix to organize category predictions against reference labels. The scikit-learn evaluation guide documents common metrics. Choose measures because they answer your question, not because a high number looks reassuring.

Let errors choose the next experiment

Inspect incorrect development predictions and group them: missing readings, unusual weather, one device type, or examples near the threshold. A repeated pattern points toward a hypothesis. If errors cluster around missing inputs, a larger network may not address the cause.

Change one important factor at a time where practical. Add a feature, correct a label, adjust a threshold, or compare a model family. Record the hypothesis, change, dataset version, settings, metrics, and examples that improved or worsened.

Use validation results for these decisions. Repeatedly checking the final test set while experimenting makes it part of the selection process. When development is finished, evaluate the selected approach once on an appropriate untouched test, and report the result without hiding unfavorable cases.

Plan what happens after the notebook

Inputs can change after deployment: sensors drift, schedules change, or the user population differs. Monitor missing values, input ranges, prediction patterns, and verified outcomes where available. Define when to investigate and when to stop using the model.

A useful project report includes the problem, data limits, baseline, selection process, test results, and operating boundaries. State untested conditions explicitly. Saving a trained model is only one deliverable; someone else also needs enough information to reproduce and interpret it.

For hardware, retain explicit safety behavior outside the model. The next stage of this publication combines AI with control, where an uncertain estimate must not become an unrestricted command.

Important terms

Precision
The fraction of predicted positives that are true positives.
Recall
The fraction of actual positives that are found.
Confusion matrix
A table comparing predicted and reference categories.
Error analysis
Inspecting failures to identify recurring causes.
Experiment log
A record of hypotheses, changes, settings, and results.
Monitoring
Checking ongoing system inputs, outputs, and outcomes.

Mini project: Write an improvement brief

  1. Choose your ML04 project or another small task.
  2. State the user, target, input availability, and baseline.
  3. List five observed or plausible errors and group them by cause.
  4. Select one change with a testable hypothesis and validation measure.
  5. Finish with a final-test plan and conditions where the result should not be used.

Common mistakes and debugging

  • Optimizing accuracy on a rare-event task: inspect false negatives and false positives.
  • Changing everything at once: preserve enough control to interpret results.
  • Selecting with the test set: reserve it for final assessment.
  • Ignoring behavior after training: specify missing-input and monitoring responses.

Independent challenge

For the five-positive-days example, propose a result with lower accuracy but higher recall. Explain the tradeoff without claiming one metric is universally best.

Check your understanding: 10 questions

  1. Why define the user action first?

  2. What does precision measure?

  3. Why can 95 percent accuracy be useless in the example?

  4. Which data guides experiments?

  5. Why monitor a deployed model?

  6. In your own words, what does “Precision” mean?

  7. In your own words, what does “Recall” mean?

  8. In your own words, what does “Confusion matrix” mean?

  9. In your own words, what does “Error analysis” mean?

  10. In your own words, what does “Experiment log” mean?

Quiz answers

Reveal all 10 answers after your attempt
  1. It determines which errors and operating constraints matter.
  2. How many predicted positives are actually positive.
  3. Always predicting negative misses all five positive days.
  4. Development and validation data, not the reserved final test.
  5. Input conditions and outcomes can change after development.
  6. The fraction of predicted positives that are true positives.
  7. The fraction of actual positives that are found.
  8. A table comparing predicted and reference categories.
  9. Inspecting failures to identify recurring causes.
  10. A record of hypotheses, changes, settings, and results.

Summary

Improve ML through clearer targets, legitimate baselines, suitable metrics, and error-driven experiments. Preserve final testing and document the conditions your evidence actually covers.

Continue learning

ML10 organizes these skills into a longer route from beginner experiments to advanced system design.

Choose a connected learning path

Sources and further reading

Prepared 2026-09-18. Draft — metric definitions and illustrative arithmetic checked