Purlin is a communication framework for distributed GPU inference that separates collective communication into three layers: semantics specification, a shared orchestration protocol (SNAC), and hardware-specific datapaths (Atom). Evaluated on A100, H200, and B200 GPUs, Purlin achieves significant speedups in latency and bandwidth, with throughput improvements of up to 1.37x for LLM serving and 2.85x for online inference.