typed-lm is an open-source Rust framework that converts large language models like Llama and Qwen into typed semantic-routing APIs, enabling deterministic inference with millisecond latency by returning structured outputs (booleans, choices, scores) instead of generated text. It includes a trainer for adapter-based specialization and supports quantization for efficient deployment on GPU and CPU.