- Represent inputs as rows of features.
- Split examples before fitting.
- Train and use LinearRegression.
- Compare model MAE with a mean baseline.
- Explain why synthetic data cannot establish real robot performance.
Before you begin
Understand variables, lists, functions, and virtual environments. If following publication dates, use PY08 as a companion prerequisite for installation before running this lesson.
Define a deliberately small experiment
We will predict travel time in seconds from distance in meters. The data is invented for teaching, not collected from a physical cart. It follows a roughly linear relationship with small variations so that you can inspect every row and understand the fitting process.
The library scikit-learn supplies standard model fitting and evaluation tools. We import LinearRegression for a straight-line model, train_test_split for a repeatable split, and mean_absolute_error for an error measured in seconds. Its official getting-started guide explains the common fit and predict interface. scikit-learn getting started.
This project uses an ordinary computer. No GPU, camera, or robot is required. Practice time includes environment setup and experiments, which is separate from reading time.
Install in a project environment
Use a supported Python 3 installation and create a project folder. In a terminal opened there, create an environment and install scikit-learn using the environment's interpreter. The commands below target macOS or Linux. On Windows, create with py -m venv .venv and use .venv\Scripts\python.exe in place of .venv/bin/python.
Using the interpreter's full path avoids confusion about which environment receives the package. If installation fails because the Python release is unsupported, consult the current scikit-learn installation requirements rather than forcing incompatible packages.
python3 -m venv .venv
.venv/bin/python -m pip install scikit-learn
.venv/bin/python cart_model.pyThe first command makes an isolated environment. The second installs the library there. Save the next example as cart_model.py before running the third command.
Expected result: Installation messages first; then the script prints model and baseline errors plus a prediction.
Fit, compare, and predict
Copy this complete Python example. Each inner input list is one row with one feature. The target list contains the corresponding observed travel times. Keeping that alignment is essential: a distance paired with the wrong time teaches the wrong relationship.
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error
from sklearn.model_selection import train_test_split
distances = [[1], [2], [3], [4], [5], [6], [7], [8], [9], [10], [11], [12]]
times = [5, 7, 10, 11, 13, 16, 17, 19, 22, 23, 25, 28]
x_train, x_test, y_train, y_test = train_test_split(
distances, times, test_size=0.25, random_state=7
)
model = LinearRegression()
model.fit(x_train, y_train)
predictions = model.predict(x_test)
training_mean = sum(y_train) / len(y_train)
baseline = [training_mean] * len(y_test)
print(f"Model MAE: {mean_absolute_error(y_test, predictions):.2f} s")
print(f"Baseline MAE: {mean_absolute_error(y_test, baseline):.2f} s")
print(f"At 6.5 m: {model.predict([[6.5]])[0]:.2f} s")The split reserves three rows for testing and nine for fitting. random_state fixes this demonstration's split. fit learns the line only from training rows. predict evaluates held-out inputs. The baseline always predicts the training mean, avoiding information from test labels. The final double brackets describe one new row with one feature; [0] retrieves its first prediction.
Expected result: Three lines report model MAE in seconds, baseline MAE in seconds, and a prediction near 16–17 seconds for 6.5 meters. Inspect exact values from your run; this draft's scikit-learn example has not yet been executed in the authoring environment.
Interpret the errors
MAE averages absolute differences between predicted and actual times. An MAE of 0.5 seconds would mean the average absolute error across those particular test rows is half a second. It does not mean every prediction is within half a second, and it is not a percentage accuracy.
The learned line should suit this constructed dataset better than ignoring distance. If it does not, inspect data alignment, changes to values, and the split. Record each test prediction and error rather than relying only on the mean.
Twelve synthetic observations and one split are insufficient evidence for real deployment. We deliberately chose a pattern favorable to a line. A genuine cart dataset needs varied conditions and an evaluation split matching intended use, such as unseen trips rather than near-duplicate readings.
Modify one thing at a time
Change a single target value and rerun. Observe how the line and errors respond. Then restore it and try another split seed. Variation shows how much a tiny experiment can depend on a few examples.
Do not choose the seed that makes the score look best and report it as an independent result. That uses the test as a selection tool. ML05 explains how validation and final testing keep such development choices separate.
If a ModuleNotFoundError appears, run the file with the same interpreter used for installation. If a two-dimensional input error appears, restore the inner brackets. If the file cannot be found, check the folder and filename rather than reinstalling the library.
Important terms
- Regression
- Predicting a numeric quantity.
- Feature matrix
- Rows of examples and columns of inputs.
- Fit
- Learn model parameters from training examples.
- MAE
- Mean absolute error, in the target's units.
- Random seed
- A setting making a randomized operation repeatable.
- Synthetic data
- Data constructed rather than measured from the intended environment.
Mini project: Record a reproducible first model
- Create the environment and save the complete script.
- Record Python and scikit-learn versions from your environment.
- Run the script and preserve all three output lines.
- Print each held-out prediction beside its actual value and inspect the largest error.
- Finish with a short note explaining the baseline, split, and synthetic-data limitation.
Common mistakes and debugging
- Using a flat feature list: use one inner row per example.
- Fitting on all rows before splitting: hold out test rows first.
- Calculating the baseline from test labels: use training information only.
- Calling MAE accuracy: report the unit and meaning of the error.
Independent challenge
Add a second feature representing load. Before editing the code, explain how every row and the new-input prediction must change to keep two consistent columns.
Check your understanding: 10 questions
Why are distances nested in lists?
How many rows are held out?
What does fit change?
Why calculate the baseline from training labels?
Does a small synthetic-data MAE prove robot reliability?
In your own words, what does “Regression” mean?
In your own words, what does “Feature matrix” mean?
In your own words, what does “Fit” mean?
In your own words, what does “MAE” mean?
In your own words, what does “Random seed” mean?
Quiz answers
Reveal all 10 answers after your attempt
- scikit-learn expects rows of examples and columns of features.
- Three, because 25 percent of twelve is three.
- It learns the regression model's parameters from the training data.
- Test labels must not supply information used to construct predictions.
- No. It demonstrates behavior on this constructed dataset and split only.
- Predicting a numeric quantity.
- Rows of examples and columns of inputs.
- Learn model parameters from training examples.
- Mean absolute error, in the target's units.
- A setting making a randomized operation repeatable.
Summary
The complete workflow is prepare, split, fit, predict, and compare. A small runnable experiment is valuable when its data assumptions and evaluation limits remain explicit.
Continue learning
ML05 explains data quality, leakage, and splits so your next experiment measures something meaningful.
- Python Libraries, pip, and Virtual Environments
- Working With Data Using NumPy and pandas
- Training Data: Quality, Bias, Cleaning, and Data Splits
- Build a Python Sensor-Data Logger and Analyzer
Sources and further reading
Prepared 2026-09-18. Draft — current API documentation checked; scikit-learn execution pending (library unavailable in authoring environment)