Technical documentation tracing a single token through prefill inference on a 32-GPU MoE model using six parallelism techniques: tensor, context, sequence, expert, data, and pipeline parallelism. The article maps communication patterns and GPU placement across two mesh configurations for attention and expert operations.