Phonon-2 is an open-source speech recognition model that achieves top accuracy at 164 MB download size, matching larger models on standard benchmarks while exceeding performance on parliamentary and meeting speech. It runs efficiently across Mac, Linux, Windows, and NVIDIA GPUs, converting an hour of audio to text in roughly 20 seconds on a MacBook Air.