VibeSys, a multi-agent system, built a serving engine for Qwen3.5-397B on AMD MI300As that achieves 2,242 tok/s of goodput, 2.33× better than tuned SGLang. Over 105 hours, agents iteratively optimized the engine for the hybrid mixture-of-experts model on a multi-turn chat workload without human code writing.
A technical article on Kolibri, Aleph Alpha's sovereign language model, examining architectural trade-offs in autoregressive models including parameter allocation, FLOPs scaling, and sequence-mixer state. The piece provides an interactive tool for configuring and comparing model architectures against recent open-weight releases.
Technical documentation tracing a single token through prefill inference on a 32-GPU MoE model using six parallelism techniques: tensor, context, sequence, expert, data, and pipeline parallelism. The article maps communication patterns and GPU placement across two mesh configurations for attention and expert operations.
lily-qwen3.8-flash-next is a Metal inference server for Apple Silicon that serves the Qwen3.8-Flash-Next model with hand-optimized kernels and speculative decoding, achieving 2.7–4.2× faster prefill and 2.1–3.6× faster decode than comparable systems. It requires an M5 GPU or newer, macOS 26, and 64–128 GB of unified memory, and provides an OpenAI-compatible API with thinking enabled by default.