What you will learn
  • Distinguish local inference from model training.
  • Run a small classifier without cloud requests.
  • Keep training and test examples separate.
  • Interpret accuracy alongside a confusion matrix.
  • Measure prediction time without overstating results.

Before you begin

Understand features, labels and held-out data. Use Raspberry Pi OS on a Pi 4 or 5; this exercise does not need GPIO or a camera.

Local AI does not have to mean a chatbot

AI includes much smaller systems than language models. A decision tree can classify a short vector of measurements using a sequence of learned comparisons. It is an excellent first local model because the data fits in memory and the prediction is easy to inspect.

Local inference means the input is processed on your own computer rather than sent to a remote inference service. Training chooses model parameters from examples; inference applies the trained model. Both happen on the Pi here, but a larger project may train elsewhere and transfer a trusted model for local inference.

Our example uses the bundled Iris dataset: measurements of flowers and three class labels. It is an educational dataset, not a camera detector or a universal plant-identification system. The point is to exercise the whole learning workflow with inputs whose meaning you can explain.

Prepare one known environment

Use a Pi 4 or 5 with storage, cooling appropriate to the board, and a suitable power supply. No accelerator is needed for this tiny classifier. On Raspberry Pi OS run sudo apt update and sudo apt install python3-sklearn. Then run the script with the system python3 that sees the installed package.

If you choose pip instead, make a virtual environment and install packages there; do not mix installation approaches blindly or override an externally managed system environment. Record Python, OS and scikit-learn versions. Distribution packages can lag the latest online documentation, so check API compatibility rather than insisting every computer has identical version numbers.

The dataset loads from scikit-learn’s installed resources. The script makes no cloud inference requests and needs no API key. Installation may require a network connection; running a downloaded application locally is different from never using the internet at any stage.

RequirementRole
Pi 4 or 5 with Raspberry Pi OSRuns the model and evaluation
Working Python and packaged scikit-learnProvides the ML implementation
Terminal access and project folderRuns and stores the script
No camera, sensor or cloud keyData comes from the bundled teaching dataset

A complete train-test experiment

Save the program as local_classifier.py and run python3 local_classifier.py. Each feature row contains four measurements, not four unrelated training examples. The target contains class identifiers for those rows.

The split holds back a quarter of the examples. stratify preserves representation of each class in this small split, while random_state makes the partition repeatable. A fixed seed is useful for debugging; it does not make one partition a definitive performance estimate.

The tree’s depth limit bounds its complexity. fit sees training data only. predict sees held-out feature rows, after which metrics compare predictions with held-out labels. The timing covers one prediction call, excluding loading and training; it is a local observation, not a claim about another Pi.

Python
from time import perf_counter
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score, confusion_matrix

data = load_iris()
train_x, test_x, train_y, test_y = train_test_split(
    data.data, data.target, test_size=0.25,
    random_state=7, stratify=data.target
)
model = DecisionTreeClassifier(max_depth=3, random_state=7)
model.fit(train_x, train_y)
started = perf_counter()
predicted = model.predict(test_x)
elapsed = perf_counter() - started
print("Features:", list(data.feature_names))
print("Classes:", list(data.target_names))
print("Accuracy:", round(accuracy_score(test_y, predicted), 3))
print(confusion_matrix(test_y, predicted, labels=[0, 1, 2]))
print(f"Batch prediction: {elapsed * 1000:.3f} ms")

The held-out labels are used only for scoring. The confusion matrix fixes the class order so rows (true classes) and columns (predictions) can be interpreted. perf_counter is a monotonic high-resolution elapsed-time clock.

Expected result: Feature and class names, an accuracy between zero and one, a 3-by-3 confusion matrix, and a measured duration. Exact duration depends on your system; no benchmark number is promised.

Read errors, not just a score

Suppose a confusion matrix has mostly counts on its diagonal. That means many predictions match their true classes. Off-diagonal cells reveal which classes are being confused. A single accuracy number hides that pattern, so use both views.

The dataset is small. Changing the random seed changes which examples are held out and may change the score. Repeatedly choosing the seed that gives the best score turns the test set into a selection tool. For serious model selection, use separate validation procedures and reserve a final test set.

Try a shallow tree and a deeper tree as a learning experiment. Explain why extra flexibility can fit training examples better without improving unfamiliar inputs. Report the experiment honestly instead of renaming the winning result “the model’s accuracy.”

Before a real deployment, test the kind of data the device will encounter. A model trained on measured flower dimensions cannot accept raw camera pixels without a different input pipeline. Similar-looking arrays can represent completely different quantities.

Decide whether a larger model belongs on your Pi

The next model may require far more memory, storage and computation. Check supported input sizes, runtime, architecture, licensing and dependencies before downloading it. Disk size alone is not a prediction of peak RAM use, and a model that loads can still be too slow for the task.

Quantization reduces the numerical precision used to store or compute parameters. It can save resources, but the effect on quality and speed depends on the model and runtime. An accelerator only helps operations it supports; it is not a universal speed button.

Measure cold start separately from repeated prediction, and measure the complete pipeline when using sensors or cameras. Image capture, resizing and output handling can dominate time. Keep the Pi’s power and temperature stable when comparing experiments.

Local processing gives you control over where inference runs, but a local application can still write sensitive logs or include network features. Review what you actually execute. Never load an untrusted Python pickle or joblib artifact just because it is described as model weights; such formats can execute code during loading.

Important terms

Inference
Applying a trained model to an input.
Decision tree
A model that routes inputs through learned feature comparisons.
Held-out set
Examples kept out of fitting for evaluation.
Confusion matrix
Counts of true versus predicted class labels.
Quantization
Representing model values with reduced numerical precision.

Mini project: A reproducible local result

  1. Record OS, Python and library versions.
  2. Run the script and label the rows and columns of its matrix.
  3. Change only max_depth and compare the result.
  4. Write a short report separating observed accuracy, the test split and prediction duration. State that this is an educational dataset.

Common mistakes and debugging

  • Testing on training examples: score held-out rows instead.
  • Calling one fast prediction a device-wide benchmark: record the measured operation and environment.
  • Passing camera pixels to a four-feature classifier: respect the trained input meaning and shape.
  • Treating local files as automatically safe: obtain executable model artifacts from a trusted source.

Independent challenge

Predict what happens if the order of two feature columns is swapped only during testing. Run the experiment on a separate copy and explain why matching input semantics matters.

Check your understanding: 10 questions

  1. Does local AI require a language model?

  2. Which step learns the tree?

  3. Which step produces held-out predictions?

  4. Why keep test labels away from fitting?

  5. What does stratify help preserve?

  6. What do diagonal confusion-matrix cells count?

  7. Why is the timed batch not a cold-start measurement?

  8. Will a deeper tree always generalize better?

  9. Does this model classify a photograph?

  10. Why avoid untrusted pickle model files?

Quiz answers

Reveal all 10 answers after your attempt
  1. No. Small classifiers are also local AI systems.
  2. fit on training features and labels.
  3. predict on test features.
  4. To assess predictions on examples not used to learn parameters.
  5. Representation of each target class in the split.
  6. Correct predictions for the corresponding classes.
  7. The timer starts after data loading and model fitting.
  8. No. It can learn unnecessary detail and overfit.
  9. No. Its input is four numeric flower measurements.
  10. Loading them can execute code, not merely read inert weights.

Summary

You have a genuine local training and inference workflow with a held-out evaluation. Its usefulness comes from matching the problem and measuring honestly, not from the size of the model.

Continue learning

PI10 connects a small learned decision to a low-risk physical indicator, using the GPIO and sensor skills already practiced.

Choose a connected learning path

Sources and further reading

Prepared 2026-09-19. Editorial draft. Primary documentation checked 19 September 2026. Code requires the stated Pi environment; no physical wiring, camera or performance test is claimed.