source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
SATURDAY, OCTOBER 10, 2026
  1. 001Hacker NewsOCT · 10English

    VibeSys Builds a Qwen3.5-397B Engine on MI300A, 2.3× Tuned SGLang

    VibeSys, a multi-agent system, built a serving engine for Qwen3.5-397B on AMD MI300As that achieves 2,242 tok/s of goodput, 2.33× better than tuned SGLang. Over 105 hours, agents iteratively optimized the engine for the hybrid mixture-of-experts model on a multi-turn chat workload without human code writing.

    By matt_d
  2. 002Hacker NewsOCT · 10English

    Designing Kolibri: Architecture Trade-Offs from First Principles

    A technical article on Kolibri, Aleph Alpha's sovereign language model, examining architectural trade-offs in autoregressive models including parameter allocation, FLOPs scaling, and sequence-mixer state. The piece provides an interactive tool for configuring and comparing model architectures against recent open-weight releases.

    By frostbyte7
  3. 003Hacker NewsOCT · 09English

    Trace a Token Through TP/CP/SP/EP/DP/PP in Moe Inference

    Technical documentation tracing a single token through prefill inference on a 32-GPU MoE model using six parallelism techniques: tensor, context, sequence, expert, data, and pipeline parallelism. The article maps communication patterns and GPU placement across two mesh configurations for attention and expert operations.

    By Charles Xu
  4. 004Hacker NewsOCT · 08English

    Lily-Qwen3.8-Flash-Next

    lily-qwen3.8-flash-next is a Metal inference server for Apple Silicon that serves the Qwen3.8-Flash-Next model with hand-optimized kernels and speculative decoding, achieving 2.7–4.2× faster prefill and 2.1–3.6× faster decode than comparable systems. It requires an M5 GPU or newer, macOS 26, and 64–128 GB of unified memory, and provides an OpenAI-compatible API with thinking enabled by default.

    By kerenskiy