Laya is an open-weight model optimized for Apple Silicon that runs locally on macOS with Core ML, achieving 49-50 decisions per second in a Snake game demo with no generated tokens. The multilingual variant processes decisions in ~5ms on M3 Max with 2.78× better energy efficiency than MLX, and requires no PyTorch or external dependencies for inference.
Laya MLX is an open-weight decision model running natively on Apple Silicon, delivering typed decisions in 13.4 ms median latency with zero output tokens. It enables local inference without external dependencies, demonstrated through a Snake game where every move triggers real-time decision-making with safety layer corrections.