- Describe an image as numerical pixel data.
- Distinguish four common vision tasks.
- Explain preprocessing and model output.
- Separate OpenCV utilities from trained model capabilities.
- Design varied tests and inspect false detections.
Before you begin
Understand datasets, generalization, and basic neural-network transformations.
Images begin as measurements
A camera photograph looks like objects to a person. Software receives numerical pixel values arranged in rows and columns, often with several color channels. A vision system must transform those measurements into the information its task requires.
Pixels depend on lighting, exposure, lens, viewpoint, and motion. The same object can therefore produce very different arrays. A model should not be described as seeing exactly as a human does; we can instead measure its ability to identify relevant patterns under stated conditions.
Resolution is the image's dimensions in pixels. Higher resolution can preserve detail but increases processing and memory demands. Resizing may help performance while removing small objects, so input size is a task decision rather than an automatic quality improvement.
Choose the correct vision task
Classification assigns a label to an image or selected region, such as contains a cup. Detection identifies object instances and locations, often using bounding boxes. Segmentation assigns labels at pixel level. Tracking associates observations across frames so an object can retain an identity over time.
A detector does not necessarily track; two boxes in consecutive frames are not automatically known to be the same object. A tracker can lose identity when objects overlap or leave the view. Likewise, recognizing a cup does not reveal its exact distance from a single ordinary image without additional assumptions or information.
For counting classroom bins, detection may be appropriate. For deciding which pixels belong to a painted path, segmentation or carefully designed conventional processing may fit. Ask what information the controller needs before selecting the model.
| Task | Typical output | Question answered |
|---|---|---|
| Classification | Category scores | What is in this image? |
| Detection | Boxes and categories | Where are the objects? |
| Segmentation | Pixel labels | Which pixels belong together? |
| Tracking | Identities over time | Is this the same moving object? |
Follow a camera pipeline
A typical pipeline captures a frame, prepares it, applies a model or algorithm, interprets outputs, and displays or uses the result. Preparation may resize the frame and convert color ordering or numeric ranges. Those details must match the model's documented expectations.
OpenCV is a computer-vision library that provides image operations, camera access, geometry, and other utilities. It does not automatically supply a correct detector for every object. A modern model often needs separately obtained weights, labels, and preprocessing rules. OpenCV introduction.
For a beginner lesson, first display a captured image and verify orientation and color. Then run inference on one saved frame. Only after that should you add continuous capture. This isolates camera problems from model problems.
Camera → frame preparation → model → boxes/scores → decision → display or controllerDesign a cup-detection experiment
Specify one small goal: mark visible cups on a desk in a controlled demonstration. Choose ten consented or object-only images with different cup positions and lighting. Include images without cups and images containing similar objects. Annotate the cups yourself so there is a reference for comparison.
A false positive is a reported cup where no cup exists. A false negative is a missed cup. Record both. Raising a confidence threshold may remove false positives while increasing misses, so a larger threshold is not universally better.
For detection, location also matters. A box spanning half the desk is not necessarily a useful detection even if its category is correct. Overlap measures such as intersection over union compare predicted and reference regions. More formal evaluation belongs in later project work, but this distinction should guide your first checks.
Connect perception to decisions carefully
A display can show uncertain detections while a person watches. A robot needs explicit handling for stale frames, missing detections, and contradictory measurements. Image recognition alone should not be the only protection against collisions.
Measure end-to-end delay, not only model execution time. Capture, resizing, inference, and communication all consume time. A correct result that arrives after an obstacle has moved may be unsuitable for control. The AI + Raspberry Pi series builds on these distinctions when cameras become part of working systems.
Important terms
- Pixel
- An image sample at a row and column.
- Channel
- One component of pixel data, such as red intensity.
- Bounding box
- A rectangle locating an object candidate.
- Segmentation
- Assigning labels at pixel level.
- Tracking
- Associating objects across frames.
- False positive
- A reported detection without the corresponding target.
- Latency
- Delay between input and usable result.
Mini project: Create a vision test sheet
- Collect ten object-only desk images with varied conditions.
- Mark actual cup positions and include at least two images without cups.
- Write which output your task needs: label, box, pixels, or identity over time.
- List three failure conditions and how to record them.
- Finish with a test sheet that could evaluate any candidate detector consistently.
Common mistakes and debugging
- Treating classification as localization: choose an output that answers the task.
- Ignoring color format and resizing rules: match model documentation.
- Calling detection tracking: add temporal association when identity matters.
- Reporting only attractive examples: include absent targets and difficult conditions.
Independent challenge
Change the task from marking cups to following one selected cup as it moves. Explain which additional information and tests are required.
Check your understanding: 10 questions
What does software receive from an image?
Which task outputs pixel labels?
Does OpenCV guarantee an object model?
What is a false negative?
Why measure total latency?
In your own words, what does “Pixel” mean?
In your own words, what does “Channel” mean?
In your own words, what does “Bounding box” mean?
In your own words, what does “Segmentation” mean?
In your own words, what does “Tracking” mean?
Quiz answers
Reveal all 10 answers after your attempt
- An array of numerical pixel values, often with multiple channels.
- Segmentation.
- No. Its utilities and any chosen trained model are separate components.
- A target present in the reference that the system misses.
- Capture, preprocessing, inference, and communication all affect when a decision becomes usable.
- An image sample at a row and column.
- One component of pixel data, such as red intensity.
- A rectangle locating an object candidate.
- Assigning labels at pixel level.
- Associating objects across frames.
Summary
Vision tasks differ in their required outputs. Prepare images correctly, evaluate varied conditions, and separate perception results from the decisions that use them.
Continue learning
ML08 applies similar representation and evaluation ideas to language tasks.
- Raspberry Pi Cameras and Computer Vision
- Build a Raspberry Pi Camera Object Detector
- Robot Cameras and Computer Vision
- Computer Vision vs Image Generation
Sources and further reading
Prepared 2026-09-18. Draft — OpenCV role source checked; no camera or model execution claimed