Purlin is a communication framework for distributed GPU inference that separates collective communication into three layers: semantics specification, a shared orchestration protocol (SNAC), and hardware-specific datapaths (Atom). Evaluated on A100, H200, and B200 GPUs, Purlin achieves significant speedups in latency and bandwidth, with throughput improvements of up to 1.37x for LLM serving and 2.85x for online inference.
A developer built Sangama, an open-source project enabling distributed inference of large language models across multiple computers with limited GPU memory. Running Qwen3.5-397B across 20 GPUs with 16GB each achieved 5.5 tokens per second, with network latency rather than compute power being the primary bottleneck.