What you will learn
  • Distinguish learning setups by feedback.
  • Choose classification or regression for a target.
  • Explain what clusters do and do not mean.
  • Identify states, actions, and rewards.
  • Recognize self-supervised targets.

Before you begin

Be able to identify features, labels, and predictions.

Start with the feedback available

A school garden offers several ML tasks. Predict tomorrow's soil moisture, group plants with similar measurements, or learn an irrigation policy in a simulator. All involve data, but their feedback differs. Choosing the learning setup begins with the question and evidence available, not a favorite algorithm.

Supervised learning uses input examples paired with target outputs. Unsupervised learning examines structure without those task-specific targets. Reinforcement learning concerns actions, changing situations, and reward feedback over time. One project can combine methods; these categories need not describe an entire product exclusively.

Supervised learning: categories and quantities

For leaf images labeled damaged or undamaged, the task is classification. The target is a category. A model may output scores or probabilities before a decision rule selects a category. Binary classification has two classes; multiclass classification has more than two mutually exclusive classes. Some tasks permit multiple labels per example.

Predicting soil moisture as a percentage is regression because the target is a numerical quantity. Predicting low, medium, or high moisture is classification, even though the categories are related to numerical ranges. The required output determines the formulation.

Supervised does not mean a person watches every update. It means training has target outputs against which predictions can be compared. Targets may come from instruments, existing records, or human annotation, and each source can contain errors.

Unsupervised learning: similarity without supplied answers

Suppose you measure leaf length and width but have no species labels. Clustering groups examples according to similarity in the chosen representation. You may discover groups, but the algorithm does not establish that those groups are biological species.

The result depends on features, scaling, distance measures, and algorithm settings. If one measurement ranges in thousands and another in tenths, the larger scale can dominate some distance calculations. Changing representation can change the groups.

Use clusters as a prompt for investigation: inspect examples from each group and ask whether the grouping serves a purpose. Naming clusters after seeing them is interpretation, not a newly verified ground truth. Google's introduction to ML distinguishes clustering from classification in this way.

Reinforcement learning: consequences over time

In a simulated garden, a state might include soil moisture and time. An action might request a small amount of water or wait. A reward encodes the desired result, perhaps maintaining a useful moisture range while limiting water consumption. A policy maps information about the state to actions.

Actions affect later states. Watering now may help later, while repeatedly watering can cause harm. That sequential consequence is different from assigning labels to independent photos. Exploration means trying behavior to learn its effects; exploitation means using behavior already expected to work well.

A reward is a design choice and can be incomplete. Rewarding only plant growth could ignore water waste. This introductory example belongs in simulation, not unsupervised real watering equipment. Practical control requires constraints and domain knowledge beyond a reward formula.

Self-supervision and choosing a first task

Self-supervised learning creates targets from data itself. A text model might predict withheld words; an image method may learn from transformed views. These tasks reduce the need for manually labeled examples while retaining an explicit learning objective.

For a first project, choose a small supervised task with a clear target and easy evaluation. Predicting a numerical measurement or classifying familiar categories makes the feedback visible. Use clustering when exploring structure is itself useful, and study sequential decision-making before tackling reinforcement learning.

Write the task in one sentence: Given these inputs, produce this output, evaluated by this measure. If you cannot finish that sentence, refine the problem before choosing software.

Important terms

Classification
Predicting a category or category scores.
Regression
Predicting a numeric quantity.
Clustering
Grouping examples by a chosen similarity measure.
State
Information describing the current situation.
Action
A choice affecting the environment.
Reward
Feedback used to define an RL objective.
Policy
A rule or learned mapping for selecting actions.

Mini project: Formulate three garden tasks

  1. Write one labeled image task, one numeric prediction task, and one grouping task.
  2. Name inputs, output, and feedback for each.
  3. For a simulated watering policy, list state, actions, reward, and one ignored consequence.
  4. Finish by selecting the task with the clearest available data and evaluation.

Common mistakes and debugging

  • Calling numbered class labels regression: a category code is not necessarily a measurable quantity.
  • Treating clusters as verified natural categories: inspect and validate their meaning.
  • Calling every feedback loop reinforcement learning: an explicit controller need not learn a policy.
  • Designing rewards without considering side effects: state constraints and omitted outcomes.

Independent challenge

Reformulate moisture prediction as both regression and classification. Explain what information the category version loses.

Check your understanding: 10 questions

  1. What feedback defines supervised learning?

  2. Is predicting a temperature value classification?

  3. Do clusters prove species membership?

  4. What makes RL sequential?

  5. How does self-supervision obtain targets?

  6. In your own words, what does “Classification” mean?

  7. In your own words, what does “Regression” mean?

  8. In your own words, what does “Clustering” mean?

  9. In your own words, what does “State” mean?

  10. In your own words, what does “Action” mean?

Quiz answers

Reveal all 10 answers after your attempt
  1. Examples include target outputs for comparison.
  2. Normally regression, because temperature is a numerical quantity.
  3. No. They reflect the chosen representation and similarity method.
  4. Actions influence later states and rewards.
  5. It derives them from the data itself.
  6. Predicting a category or category scores.
  7. Predicting a numeric quantity.
  8. Grouping examples by a chosen similarity measure.
  9. Information describing the current situation.
  10. A choice affecting the environment.

Summary

Choose a learning setup from the task and available feedback. Categories, quantities, similarity groups, and action policies solve different problems.

Continue learning

ML03 opens the training loop to explain how predictions, error, and updates connect.

Choose a connected learning path

Sources and further reading

Prepared 2026-09-18. Draft — learning categories source checked; RL exercise explicitly simulated