- Locate a colored hand marker in an image.
- Normalize a pixel position across camera sizes.
- Reject missing or ambiguous visual input.
- Distinguish rule-based vision from learned hand recognition.
Before you begin
You can run the LED bridge and understand that an image is a grid of pixel values.
Make the gesture measurable
A human recognizes a wave despite different hands, sleeves and rooms. Reproducing that flexibility is a difficult perception problem. We will begin with a controlled gesture: moving a bright green marker attached to your hand from the left half of the camera image to the right half. The marker's position becomes a measurable input.
This is computer vision using programmed color rules, not a trained AI model. It provides the camera-to-hardware pipeline needed for a later learned hand recognizer, without disguising color thresholding as intelligence. The distinction matters because a rule's limitations are tied directly to its color range and scene.
Use a large green card or fabric patch, not a light shone toward someone's eyes. Position the camera on a stable surface, point it at your own desk and clear similar green objects from the background. No photos need to be stored or uploaded.
USB camera → computer: color mask → marker position
→ valid zone? → Uno LED → observe responsePrepare the computer and the LED bridge
The reference board is an Arduino Uno R3 connected by USB. Start with its built-in LED, so this output needs no external circuit. A laptop performs inference; the Arduino accepts a tiny set of commands. A Raspberry Pi can replace the laptop after its Python setup works. No cloud service is required.
Download bridge.py beside your Python script and the LED sketch. Save the sketch in a folder named led_bridge, open it in Arduino IDE, select Uno and your port, then upload. Close Serial Monitor before opening the port from Python. The kit guide supplies Windows and Raspberry Pi setup alternatives.
On macOS or Linux, create an environment with python3 -m venv .venv, activate it with source .venv/bin/activate, and install the serial library using python -m pip install pyserial. Find your port with python -m serial.tools.list_ports; replace the sample port in the code. Never install a package called bridge: that module is the downloaded file.
Commands are newline-terminated text at 115200 baud. LED_ON means illuminate the indicator; LED_OFF and STOP turn it off. The firmware acknowledges valid messages and turns the LED off after 600 milliseconds without a refreshed light command. Bridge.hold refreshes a requested action for a bounded interval and then sends STOP. An acknowledgement proves a message was parsed; looking at the board checks the physical output.
| Part | Purpose |
|---|---|
| Uno R3 and data-capable USB cable | Low-voltage output and communication |
| Computer with Python 3 | Inference and command validation |
| Built-in LED | Safe visible actuator substitute |
Understand the mask and its coordinates
Install OpenCV with python -m pip install opencv-python. Download vision.py beside bridge.py. Its marker function changes OpenCV's BGR image into HSV, which separates hue from saturation and brightness. It selects a green hue band, finds connected selected regions and uses the largest sufficiently large region as the marker.
The returned x position is divided by image width. A normalized x of0.25 means a quarter of the way from the image's left edge, whether the picture is640 or1280 pixels wide. Area is divided by the image's total pixel count. Normalization makes a threshold easier to interpret, though it does not remove changes in camera perspective or distance.
A missing region returns None rather than a fabricated position. Treat that as a stop condition. A second green object may become the largest region; the program cannot tell which green patch belongs to your hand. That is an example of a perception assumption you should write down before connecting actuators.
| Extra part | Purpose |
|---|---|
| Supported USB UVC webcam | Camera input to computer |
| Bright green card or fabric marker | Controlled hand target |
| Stable camera mount | Repeatable viewpoint |
Use a dead zone between gestures
Save this as gesture_light.py, set your actual port and run it. The image's left zone requests a short light pulse; its right zone requests off. The middle zone is deliberately neutral so small hand movements do not continually reverse a decision. The program runs for20seconds and leaves the light off.
import time
import cv2
from bridge import Bridge
from vision import marker
camera = cv2.VideoCapture(0)
if not camera.isOpened():
raise RuntimeError('USB camera could not open')
try:
with Bridge('/dev/ttyACM0') as board:
end = time.monotonic() + 20
while time.monotonic() < end:
ok, frame = camera.read()
if not ok:
raise RuntimeError('Camera frame unavailable')
result = marker(frame)
command = 'LED_OFF'
if result is not None and result[0] < 0.35:
command = 'LED_ON'
board.send(command)
print(result, command)
time.sleep(0.1)
finally:
camera.release()VideoCapture opens a USB camera; a Raspberry Pi ribbon camera needs its dedicated camera stack instead. Every frame starts with an off decision and only a valid marker in the left zone changes it. The loop refreshes the LED command, while the firmware timeout turns it off if the computer stalls. Finally releases the camera even after an error.
Expected result: The terminal prints marker coordinates or None and the chosen command. The LED lights while a valid green patch occupies the left zone and goes off when the marker moves away, disappears or the program ends.
Progress from a position rule to a learned gesture
Test left, center, right and no marker in the same lighting. Then repeat under different light and note which assumptions break. If recognition flickers, inspect the mask and adjust the scene before widening thresholds blindly. A wide green range may admit more background objects.
To add actual machine learning, first define gestures and collect labeled examples from multiple sessions. Features could include normalized hand landmarks produced by a documented pose model. Train a classifier on one set of sessions and evaluate on different sessions or people. Do not randomly split adjacent video frames and mistake near-duplicates for independent tests.
A camera command is not an emergency stop. Keep this lesson attached to an indicator. A motor version needs an independent stop control, short command leases and sensor interlocks even when the visual model seems accurate.
Important terms
- Mask
- An image marking pixels selected by a rule.
- HSV
- A color representation separating hue, saturation and value.
- Normalized coordinate
- A position expressed relative to image dimensions.
- Dead zone
- An input region that intentionally causes no active response.
- Landmark
- A consistently defined point such as a finger joint in a pose model.
Mini project: Record a gesture test matrix
- Try left, middle, right and missing-marker inputs.
- Repeat them under two lighting conditions without changing code.
- Place a second green object in view and explain any failure.
- Write a short specification for a learned recognizer that could overcome that failure.
Common mistakes and debugging
- Calling the color detector a trained model: it uses fixed rules.
- Confusing your left with the image's left: verify camera orientation before labeling actions.
- Reusing the last valid position after marker loss: treat missing input as off.
Independent challenge
Require the left-zone observation in three consecutive frames before turning the LED on. Reset the counter whenever the observation is missing or outside the zone.
Check your understanding: 10 questions
Why normalize x?
What does None mean?
Is this color mask machine learning?
Why is a middle dead zone useful?
Why avoid splitting adjacent frames across training and testing?
In your own words, what does “Mask” mean?
In your own words, what does “HSV” mean?
In your own words, what does “Normalized coordinate” mean?
In your own words, what does “Dead zone” mean?
In your own words, what does “Landmark” mean?
Quiz answers
Reveal all 10 answers after your attempt
- To express position as a fraction of image width rather than a resolution-specific pixel value.
- No sufficiently large matching region was found, so the output should remain off.
- No; its thresholds are explicitly programmed.
- It reduces unstable changes near a decision boundary.
- Nearly identical frames can leak scene information and exaggerate generalization.
- An image marking pixels selected by a rule.
- A color representation separating hue, saturation and value.
- A position expressed relative to image dimensions.
- An input region that intentionally causes no active response.
- A consistently defined point such as a finger joint in a pose model.
Summary
A controllable gesture interface begins with a measurable visual feature, explicit missing-input behavior and a bounded output. This marker experiment teaches that pipeline; a learned pose model can replace perception after you establish an honest evaluation method.
Continue learning
The next project introduces a genuinely trained image detector and connects its result to the same Arduino output boundary.
- Connect AI Object Detection to Arduino Outputs
- Robot Cameras and Computer Vision
- Computer Vision: How Machines Understand Images
Sources and further reading
Prepared 2026-09-18. Editorial draft. Primary documentation checked; hardware, camera and audio behavior still require a physical test.