A developer built Sangama, an open-source project enabling distributed inference of large language models across multiple computers with limited GPU memory. Running Qwen3.5-397B across 20 GPUs with 16GB each achieved 5.5 tokens per second, with network latency rather than compute power being the primary bottleneck.