What you will learn
  • Use a trained single-class object detector.
  • Interpret a bounding box without confusing it with identity.
  • Map a detection result to an allowed Arduino action.
  • Evaluate false positives and missed detections.

Before you begin

Complete the camera and serial exercises and understand image coordinates.

Detection is a location estimate

Image classification assigns a label to an image. Object detection also estimates where instances of a class appear, usually as rectangles called bounding boxes. A frame can contain several detected objects, and a rectangle can still be wrong. Detection is not proof that the model understood a scene.

Our baseline is OpenCV's built-in HOG/SVM pedestrian detector. HOG summarizes local edge directions; the support-vector-machine classifier was trained to distinguish pedestrian-shaped image windows. This is genuine machine learning, although it is older and less flexible than many modern deep detectors. Its advantage for a first build is that the weights are available through OpenCV without an extra model download.

The detector handles one class: pedestrians. It cannot identify a particular person, count everyone reliably in a crowd or recognize cups and cars. We use a still photograph first, making the exact input repeatable. Only an LED responds, so a wrong rectangle becomes a visible learning opportunity rather than a hazardous action.

Image → computer: trained pedestrian detector → boxes
      → any accepted box? → USB → Uno LED → visual check

Prepare the computer and the LED bridge

The reference board is an Arduino Uno R3 connected by USB. Start with its built-in LED, so this output needs no external circuit. A laptop performs inference; the Arduino accepts a tiny set of commands. A Raspberry Pi can replace the laptop after its Python setup works. No cloud service is required.

Download bridge.py beside your Python script and the LED sketch. Save the sketch in a folder named led_bridge, open it in Arduino IDE, select Uno and your port, then upload. Close Serial Monitor before opening the port from Python. The kit guide supplies Windows and Raspberry Pi setup alternatives.

On macOS or Linux, create an environment with python3 -m venv .venv, activate it with source .venv/bin/activate, and install the serial library using python -m pip install pyserial. Find your port with python -m serial.tools.list_ports; replace the sample port in the code. Never install a package called bridge: that module is the downloaded file.

Commands are newline-terminated text at 115200 baud. LED_ON means illuminate the indicator; LED_OFF and STOP turn it off. The firmware acknowledges valid messages and turns the LED off after 600 milliseconds without a refreshed light command. Bridge.hold refreshes a requested action for a bounded interval and then sends STOP. An acknowledgement proves a message was parsed; looking at the board checks the physical output.

PartPurpose
Uno R3 and data-capable USB cableLow-voltage output and communication
Computer with Python 3Inference and command validation
Built-in LEDSafe visible actuator substitute

Prepare an image and the vision helper

Install OpenCV in the active environment, then download vision.py beside bridge.py. Choose a photograph you own or have permission to use, showing a standing person fully in frame. Save it as scene.jpg. Include a second image with no people for a negative test. Keep both files local.

The helper checks the image, reduces very wide frames to a maximum width of640pixels and calls detectMultiScale. This examines windows at several scales because people can occupy different amounts of the picture. It filters using an example SVM score threshold and maps the rectangles back to the original image's coordinates.

A score above0.5 in this helper is not a50% probability. It is a model-specific decision margin used as a chosen threshold. Raising it may reduce false detections while increasing misses. You need examples from your own scene to judge that tradeoff. Images still need enough pixels to contain the detector's64×128 window.

Input or partPurpose
Two permitted JPG imagesRepeatable positive and negative tests
Computer with OpenCVRuns the trained detector
Uno LED bridgeBounded visible output

Draw results before trusting the output

Run this script from the folder containing scene.jpg, bridge.py and vision.py. It writes annotated.jpg so you can inspect the actual boxes. Replace the serial port. A successful run with zero boxes should leave the LED off; that is a valid result, not automatically a software error.

Python
import cv2
from bridge import Bridge
from vision import people

frame = cv2.imread('scene.jpg')
if frame is None:
    raise FileNotFoundError('Cannot read scene.jpg')
boxes = people(frame)
for x, y, width, height in boxes:
    cv2.rectangle(frame, (x, y), (x + width, y + height), (0, 255, 0), 2)
if not cv2.imwrite('annotated.jpg', frame):
    raise RuntimeError('Could not save annotated image')
print('Accepted boxes:', len(boxes))
with Bridge('/dev/ttyACM0') as board:
    board.hold('LED_ON' if boxes else 'LED_OFF', 1.0)

imread loads a BGR image, and people applies the included detector. Each rectangle uses a top-left coordinate plus width and height; adding the dimensions gives the opposite corner. imwrite saves visual evidence rather than requiring a desktop display window. Only the Boolean presence of boxes reaches the Arduino.

Expected result: A box count is printed and annotated.jpg is saved. If at least one accepted box exists, the LED lights for one second. Detection accuracy depends on the image; no particular count is guaranteed.

Build a small, honest detection report

Collect ten permitted images with a visible standing person and ten without one, including difficult backgrounds. Label them before running the detector. Count false alarms on empty images and misses on person images. Also inspect duplicate or badly positioned boxes: a yes/no score alone hides localization errors.

A live extension uses a supported USB camera and processes one fresh frame at a time. Timestamp the observation and expire decisions if the camera stops producing frames. If processing takes longer than the LED watchdog window, pulses may flicker; that is a reason to measure latency, not to remove the timeout from moving systems.

If the baseline is inadequate, keep your test set and compare a modern detector using its official model documentation and license. Change one part at a time. Retaining the same output protocol makes it possible to improve perception without rewriting the Arduino firmware.

Important terms

Bounding box
A rectangle estimating an object's image location.
HOG
Histogram of Oriented Gradients, an edge-pattern descriptor.
SVM
Support vector machine, a learned classification method.
False positive
A detection where the target class is absent.
Latency
Time between receiving an input and producing its result.

Mini project: Evaluate a detector before connecting outputs

  1. Run one positive and one negative image with the board disconnected from the script's output step.
  2. Inspect annotated boxes rather than relying only on the count.
  3. Run the20-image test and record false alarms and misses.
  4. Connect the LED bridge only after the perception outputs are understood.

Common mistakes and debugging

  • Assuming no box proves no person: detectors miss objects.
  • Treating detection as identification: a pedestrian label says nothing about who the person is.
  • Comparing models on different images: use the same labeled test set.

Independent challenge

Save a CSV with filename, expected person presence, detected presence and processing duration for every test image. Use it to compare two thresholds.

Check your understanding: 10 questions

  1. How does detection differ from classification?

  2. Does this detector recognize every object class?

  3. Is its margin a calibrated probability?

  4. Why save annotated images?

  5. What should happen if a live camera becomes stale?

  6. In your own words, what does “Bounding box” mean?

  7. In your own words, what does “HOG” mean?

  8. In your own words, what does “SVM” mean?

  9. In your own words, what does “False positive” mean?

  10. In your own words, what does “Latency” mean?

Quiz answers

Reveal all 10 answers after your attempt
  1. Detection estimates object locations as well as classes; classification typically labels the image or region.
  2. No; the included trained baseline detects pedestrian-like image windows.
  3. No. The chosen numeric threshold is not a percentage confidence.
  4. They reveal misplaced, duplicate or spurious boxes hidden by a simple count.
  5. Its last observation should expire and the system should use a defined safe output.
  6. A rectangle estimating an object's image location.
  7. Histogram of Oriented Gradients, an edge-pattern descriptor.
  8. Support vector machine, a learned classification method.
  9. A detection where the target class is absent.
  10. Time between receiving an input and producing its result.

Summary

A detector produces estimates that must be inspected and evaluated. Keeping perception on the computer and a small command set on the Arduino makes the system understandable and allows you to replace the model without changing its output contract.

Continue learning

Continue with an alert prototype that combines sensor evidence and vision while keeping privacy and false alarms visible.

Choose a connected learning path

Sources and further reading

Prepared 2026-09-18. Editorial draft. Primary documentation checked; hardware, camera and audio behavior still require a physical test.