¡Oído! is what cooks call out in a Spanish kitchen to confirm an order: heard, got it.

Speech-to-text for any English sentence, running entirely on an ESP32-S3 (240 MHz dual-core Xtensa LX7, 8 MB PSRAM, 16 MB flash). No cloud, no command list, no neural accelerator. Built by Lokutor. Model on Hugging Face: lokutor-ai/oido-ctc-small-int8.

Status (30 September 2026). Every transcript below comes from the exact arithmetic of the on-chip engine: the host build is bit-identical to the firmware, and firmware transcripts under Espressif's QEMU emulator match it word for word. Real-time speed is estimated from exact emulator instruction counts. Measurements on physical boards follow in the next days and will be added here.

Word error rate (%) on LibriSpeech, same text normalization for every system.

- On-chip rows use the full test sets. Laptop baselines use 500 evenly spaced utterances per set. MultiNet7 figures are from Espressif's ESP-SR benchmark page.

- The int8 engine is within 0.1 points of full precision: 3.70 / 8.23 on chip vs 3.68 / 8.11 for the original fp32 model.

- To our knowledge this is the most accurate LibriSpeech result published for any microcontroller. It is not the first open-vocabulary recognizer on one (Arm has shown Conformer models on Cortex-M55 + Ethos-U NPUs).

Robustness (300 LibriSpeech utterances under real DEMAND noise, babble and room reverb; eval/make_robust.py,

full numbers in results/robustness.json):

For the transducer, car and kitchen noise at 5 dB SNR cost under 1 point, and living-room noise about 1.7. Four-talker babble at 5 dB and very reverberant rooms are the hard cases.

- Front end: log-mel features, then 2× (3×3, stride 2) convolution subsampling to 25 Hz.

- Encoder: 16 Conformer layers (d = 176, 4 heads, relative-position attention, conv kernel 31).

- Decoding: CTC over 1024 BPE tokens. The engine also supports an RNN-T head (LSTM 320 + joint network) and a GRU language model with CTC prefix beam search.

The engine (esp32/components/tinyasr) is new C written for the ESP32-S3's PIE vector unit:

- int8 matrix kernels on EE.VMULAS.S8.ACCX(16 MACs per instruction), plus int4 outer-product kernels;

- int8 relative-position attention with a lookup-table softmax;

- dual-core scheduling;

- tiling so that each weight is streamed from flash once per 64 frames (weight traffic cut from 18 to 7.5 MB/s);

- a VAD/AGC utterance segmenter;

- an optional SSD1306 OLED that shows the live transcript.

On a laptop, with the chip's exact arithmetic. Needs Python with numpy, soundfile, sentencepiece and sounddevice.

cd esp32/host && make

./tasr_cli ../../models/nemo8.tnm recording.wav # 16 kHz mono PCM16 wav

python live_demo.py # microphone -> the firmware's VAD + engine, with ESP32 time estimatesOn a board. ESP32-S3-DevKitC-1 N16R8, an INMP441 I2S microphone (SCK→GPIO4, WS→GPIO5, SD→GPIO6, L/R→GND), and optionally a 0.96" SSD1306 OLED (SDA→GPIO8, SCL→GPIO9). Needs ESP-IDF v5.5.

esp32/tools/flash.sh /dev/ttyUSB0 models/nemo8.tnm # live microphone

TASR_OLED=1 esp32/tools/flash.sh /dev/ttyUSB0 models/nemo8.tnm # + transcript on the OLED

TASR_MODE=file esp32/tools/flash.sh /dev/ttyUSB0 models/nemo8.tnm clip.wav "reference" # prints measured RTFIn the emulator (Espressif QEMU 9.x): runs the real firmware, then reports the transcript and instruction counts.

esp32/tools/emulate.sh clip.wavThe transducer model. Its weights are NVIDIA's, distributed on NGC under NVIDIA's terms, so they are not included here. You can fetch and convert them yourself:

cd train && python fetch_nemo_small.py ../models/nemo_rnnt --transducer

python export_nemo.py ../models/nemo_rnnt ../models/rnnt8.tnm 8esp32/components/tinyasr/ on-chip engine: tasr_nemo.c (Conformer CTC/RNN-T), kernels.c (PIE SIMD), tinyasr_lm.c

(GRU LM + beam search), tasr_seg.c (VAD), tinyasr.c (streaming engine)

esp32/firmware/ ESP-IDF app: live I2S microphone or benchmark mode, OLED, partition layouts

esp32/host/ host build of the engine: tasr_cli, live_demo.py, seg_test, eval_engine.py, benchmark.py

esp32/tools/ flash.sh, emulate.sh, run_qemu.sh, bench_latency.py, mkimages.py

train/ PyTorch port of NVIDIA's model (nemo_small.py, rnnt_small.py), exporters, GRU LM training

eval/ WER normalization, robustness benchmark builder, laptop baselines

results/ benchmark outputs behind the numbers above

models/ nemo8.tnm (int8 Conformer-CTC Small) and its tokenizer

- English only.

- Text appears after each utterance, not word by word.

- Very noisy crowds and reverberant rooms remain hard.

- Speed is estimated until board measurements are published.

- Requires an ESP32-S3 with 16 MB flash and 8 MB octal PSRAM (N16R8).

- Code is licensed under the GNU GPL v3 (LICENSE).

- For products that cannot meet GPLv3 terms (for example, consumer devices that do not allow users to install modified

firmware), Lokutor offers commercial licenses and support. See COMMERCIAL.md.

- Model weights in models/are derived from NVIDIA'sstt_en_conformer_ctc_smalland remain under CC-BY-4.0. SeeNOTICE.

Lokutor also has Spanish and other-language models, an int4 profile with more compute headroom, and an on-device TTS for the same chip. Contact us for these.