What you will learn
  • Explain which device recognizes speech.
  • Decode a local WAV recording with a speech model.
  • Accept only explicitly supported commands.
  • Verify output and communication-failure behavior.

Before you begin

You can upload a sketch, run Python in an environment and identify your serial port.

Let the computer listen and the Arduino act

Saying ‘light on’ sounds like one instruction, but the machine must solve several different problems. A microphone records pressure changes as digital samples. A speech recognizer estimates words from those samples. Your program decides whether those words match an allowed command. Only then does a microcontroller change an electrical output.

Here the AI is Vosk's trained speech-recognition model running on the computer. It is not running on an Uno R3. The Arduino executes ordinary firmware, which is useful because a light command should have a small, predictable meaning. The model does not receive unrestricted access to pins, files or the operating system.

Begin with a recording rather than a live microphone. A saved input makes failures repeatable: you can run exactly the same sound through a new version and compare the result. Live listening is a later extension after the command boundary is reliable.

WAV audio → computer: speech model → exact phrase check
         → USB serial → Uno firmware → LED → visual check

Prepare the computer and the LED bridge

The reference board is an Arduino Uno R3 connected by USB. Start with its built-in LED, so this output needs no external circuit. A laptop performs inference; the Arduino accepts a tiny set of commands. A Raspberry Pi can replace the laptop after its Python setup works. No cloud service is required.

Download bridge.py beside your Python script and the LED sketch. Save the sketch in a folder named led_bridge, open it in Arduino IDE, select Uno and your port, then upload. Close Serial Monitor before opening the port from Python. The kit guide supplies Windows and Raspberry Pi setup alternatives.

On macOS or Linux, create an environment with python3 -m venv .venv, activate it with source .venv/bin/activate, and install the serial library using python -m pip install pyserial. Find your port with python -m serial.tools.list_ports; replace the sample port in the code. Never install a package called bridge: that module is the downloaded file.

Commands are newline-terminated text at 115200 baud. LED_ON means illuminate the indicator; LED_OFF and STOP turn it off. The firmware acknowledges valid messages and turns the LED off after 600 milliseconds without a refreshed light command. Bridge.hold refreshes a requested action for a bounded interval and then sends STOP. An acknowledgement proves a message was parsed; looking at the board checks the physical output.

PartPurpose
Uno R3 and data-capable USB cableLow-voltage output and communication
Computer with Python 3Inference and command validation
Built-in LEDSafe visible actuator substitute

Prepare audio that the recognizer can read

Install Vosk in the same environment with python -m pip install vosk. Download the small English US model from the official model catalog, unpack it and name its directory model. Check the model license before redistributing it. The model is a separate download from the Python library.

Record yourself saying ‘light on’ in a quiet room and export command.wav as uncompressed 16-bit PCM, mono, 16,000 samples per second. An MP3 renamed to WAV is still an MP3. Use your audio editor's export controls; the following program rejects the wrong format before inference. Add several seconds of quiet at the end if the recognizer cuts off your last word.

Recognition is uncertain. A limited vocabulary improves focus but can also encourage an incorrect match to one of the permitted phrases. The unknown token lets the recognizer represent other speech; it does not make voice control secure. This project controls an indicator, never a lock, heater or moving machine.

Decode, validate and send one bounded action

Save this as voice_light.py beside bridge.py, the model folder and your WAV file. Change the port string to your discovered port. Python's wave module reads the recording; json reads the recognizer's structured result. Run python voice_light.py.

Python
import json
import wave
from vosk import Model, KaldiRecognizer
from bridge import Bridge

with wave.open('command.wav', 'rb') as audio:
    assert (audio.getnchannels(), audio.getsampwidth(),
            audio.getframerate(), audio.getcomptype()) == (1, 2, 16000, 'NONE')
    recognizer = KaldiRecognizer(Model('model'), 16000,
        json.dumps(['light on', 'light off', '[unk]']))
    phrases = []
    while chunk := audio.readframes(4000):
        if recognizer.AcceptWaveform(chunk):
            phrases.append(json.loads(recognizer.Result()).get('text', ''))
    phrases.append(json.loads(recognizer.FinalResult()).get('text', ''))
    phrase = ' '.join(part for part in phrases if part).strip()
print('Recognized:', phrase)
commands = {'light on': 'LED_ON', 'light off': 'LED_OFF'}
if phrase in commands:
    with Bridge('/dev/ttyACM0') as board:
        board.hold(commands[phrase], 1.0)
else:
    print('No supported command; no action sent.')

The format assertion catches mismatched audio. AcceptWaveform feeds blocks to the recognizer, and FinalResult retrieves the final utterance for this short recording. The commands dictionary is an exact allowlist: generated text is never executed as code. Bridge opens the board, waits for its reset and acknowledges commands; hold refreshes the LED for one second and then stops.

Expected result: The program prints the recognized phrase. For ‘light on’, the built-in LED lights for about one second; unknown speech sends no action. Model loading can print additional diagnostic messages.

Test the boundary, not just a successful phrase

Make recordings of both allowed phrases, silence and unrelated speech. Save the expected outcome before running them. A false acceptance means an unrelated recording triggered an action; a false rejection means a valid phrase was missed. Record both, because ‘it worked once’ tells you little about reliability.

If the port is busy, close Serial Monitor and other Python programs. If the speech result is empty, inspect the exported audio format and listen to the file. If the printed command is right but the LED stays dark, test LED_ON in Serial Monitor with newline endings at115200baud, then close it before retrying. Keep recognition debugging separate from electrical debugging.

Important terms

Speech recognition
Estimating words from recorded sound.
PCM
Audio represented as regularly sampled numeric amplitudes.
Allowlist
The exact inputs an application permits to trigger an action.
Inference
Using a trained model on a new input.
Watchdog
A timeout that returns a device to a defined safe state.

Mini project: Build a four-recording test set

  1. Record light on, light off, silence and an unrelated sentence in the required format.
  2. Run each as command.wav and record expected versus actual text and LED state.
  3. Repeat one failed recording unchanged to make debugging reproducible.
  4. Confirm that stopping the host program leaves the indicator off.

Common mistakes and debugging

  • Installing Vosk without downloading its model: the library and learned parameters are separate.
  • Sending any recognized sentence to the board: use exact commands and reject everything else.
  • Treating microphone input as authentication: anyone or any speaker may produce the same phrase.

Independent challenge

Add ‘status’ as a recognized phrase that prints a local message without sending an output command. Explain why it belongs outside the hardware allowlist.

Check your understanding: 10 questions

  1. Where does the trained model run?

  2. Why start with a saved recording?

  3. What should unrelated speech do?

  4. What does the serial acknowledgement prove?

  5. Why does this demo use a one-second pulse?

  6. In your own words, what does “Speech recognition” mean?

  7. In your own words, what does “PCM” mean?

  8. In your own words, what does “Allowlist” mean?

  9. In your own words, what does “Inference” mean?

  10. In your own words, what does “Watchdog” mean?

Quiz answers

Reveal all 10 answers after your attempt
  1. On the computer or a suitably configured Pi, not on the Uno R3.
  2. It provides an identical input for repeatable comparisons and debugging.
  3. Produce no hardware action; it must not be converted into an arbitrary command.
  4. That the firmware parsed the command, not that the LED physically lit.
  5. It bounds the action and makes both activation and return to off observable.
  6. Estimating words from recorded sound.
  7. Audio represented as regularly sampled numeric amplitudes.
  8. The exact inputs an application permits to trigger an action.
  9. Using a trained model on a new input.
  10. A timeout that returns a device to a defined safe state.

Summary

Speech control is a pipeline with clear boundaries: a model estimates words, your program validates a command, and the Arduino executes a bounded action. Repeatable recordings and negative tests make the system easier to understand and improve.

Continue learning

Continue with AI-assisted smart lighting to replace spoken commands with a small model trained from examples.

Choose a connected learning path

Sources and further reading

Prepared 2026-09-18. Editorial draft. Primary documentation checked; hardware, camera and audio behavior still require a physical test.