¡Oído! is what cooks call out in a Spanish kitchen to confirm an order: heard, got it.
Speech-to-text for any English sentence, running entirely on an ESP32-S3 (240 MHz dual-core Xtensa LX7, 8 MB PSRAM, 16 MB flash). No cloud, no command list, no neural accelerator. Built by Lokutor. Model on Hugging Face: lokutor-ai/oido-ctc-small-int8.
Status (30 September 2026). Every transcript below comes from the exact arithmetic of the on-chip engine: the host build is bit-identical to the firmware, and firmware transcripts under Espressif's QEMU emulator match it word for word. Real-time speed is estimated from exact emulator instruction counts. Measurements on physical boards follow in the next days and will be added here.
Word error rate (%) on LibriSpeech, same text normalization for every system.
- On-chip rows use the full test sets. Laptop baselines use 500 evenly spaced utterances per set. MultiNet7 figures are from Espressif's ESP-SR benchmark page.
- The int8 engine is within 0.1 points of full precision: 3.70 / 8.23 on chip vs 3.68 / 8.11 for the original fp32 model.
- To our knowledge this is the most accurate LibriSpeech result published for any microcontroller. It is not the first open-vocabulary recognizer on one (Arm has shown Conformer models on Cortex-M55 + Ethos-U NPUs).
Robustness (300 LibriSpeech utterances under real DEMAND noise, babble and room reverb; eval/make_robust.py,
full numbers in results/robustness.json):
For the transducer, car and kitchen noise at 5 dB SNR cost under 1 point, and living-room noise about 1.7. Four-talker babble at 5 dB and very reverberant rooms are the hard cases.
- Front end: log-mel features, then 2× (3×3, stride 2) convolution subsampling to 25 Hz.
- Encoder: 16 Conformer layers (d = 176, 4 heads, relative-position attention, conv kernel 31).
- Decoding: CTC over 1024 BPE tokens. The engine also supports an RNN-T head (LSTM 320 + joint network) and a GRU language model with CTC prefix beam search.
The engine (esp32/components/tinyasr) is new C written for the ESP32-S3's PIE vector unit:
- int8 matrix kernels on EE.VMULAS.S8.ACCX(16 MACs per instruction), plus int4 outer-product kernels;
- int8 relative-position attention with a lookup-table softmax;
- dual-core scheduling;
- tiling so that each weight is streamed from flash once per 64 frames (weight traffic cut from 18 to 7.5 MB/s);
- a VAD/AGC utterance segmenter;
- an optional SSD1306 OLED that shows the live transcript.
On a laptop, with the chip's exact arithmetic. Needs Python with numpy, soundfile, sentencepiece and sounddevice.
cd esp32/host && make
./tasr_cli ../../models/nemo8.tnm recording.wav # 16 kHz mono PCM16 wav
python live_demo.py # microphone -> the firmware's VAD + engine, with ESP32 time estimatesOn a board. ESP32-S3-DevKitC-1 N16R8, an INMP441 I2S microphone (SCK→GPIO4, WS→GPIO5, SD→GPIO6, L/R→GND), and optionally a 0.96" SSD1306 OLED (SDA→GPIO8, SCL→GPIO9). Needs ESP-IDF v5.5.
esp32/tools/flash.sh /dev/ttyUSB0 models/nemo8.tnm # live microphone
TASR_OLED=1 esp32/tools/flash.sh /dev/ttyUSB0 models/nemo8.tnm # + transcript on the OLED
TASR_MODE=file esp32/tools/flash.sh /dev/ttyUSB0 models/nemo8.tnm clip.wav "reference" # prints measured RTFIn the emulator (Espressif QEMU 9.x): runs the real firmware, then reports the transcript and instruction counts.
esp32/tools/emulate.sh clip.wavThe transducer model. Its weights are NVIDIA's, distributed on NGC under NVIDIA's terms, so they are not included here. You can fetch and convert them yourself:
cd train && python fetch_nemo_small.py ../models/nemo_rnnt --transducer
python export_nemo.py ../models/nemo_rnnt ../models/rnnt8.tnm 8esp32/components/tinyasr/ on-chip engine: tasr_nemo.c (Conformer CTC/RNN-T), kernels.c (PIE SIMD), tinyasr_lm.c
(GRU LM + beam search), tasr_seg.c (VAD), tinyasr.c (streaming engine)
esp32/firmware/ ESP-IDF app: live I2S microphone or benchmark mode, OLED, partition layouts
esp32/host/ host build of the engine: tasr_cli, live_demo.py, seg_test, eval_engine.py, benchmark.py
esp32/tools/ flash.sh, emulate.sh, run_qemu.sh, bench_latency.py, mkimages.py
train/ PyTorch port of NVIDIA's model (nemo_small.py, rnnt_small.py), exporters, GRU LM training
eval/ WER normalization, robustness benchmark builder, laptop baselines
results/ benchmark outputs behind the numbers above
models/ nemo8.tnm (int8 Conformer-CTC Small) and its tokenizer
- English only.
- Text appears after each utterance, not word by word.
- Very noisy crowds and reverberant rooms remain hard.
- Speed is estimated until board measurements are published.
- Requires an ESP32-S3 with 16 MB flash and 8 MB octal PSRAM (N16R8).
- Code is licensed under the GNU GPL v3 (LICENSE).
- For products that cannot meet GPLv3 terms (for example, consumer devices that do not allow users to install modified
firmware), Lokutor offers commercial licenses and support. See COMMERCIAL.md.
- Model weights in models/are derived from NVIDIA'sstt_en_conformer_ctc_smalland remain under CC-BY-4.0. SeeNOTICE.
Lokutor also has Spanish and other-language models, an int4 profile with more compute headroom, and an on-device TTS for the same chip. Contact us for these.