Sference implements prefill concurrency in SGLang to maintain consistent time-to-first-token (TTFT) latency across multi-tenant GPU deployments. The approach addresses the architectural challenge of scheduling prefill work (prompt processing) versus decode work (token generation), which stress different GPU subsystems and have different queue dynamics. Prefill concurrency allows smaller requests to proceed without waiting behind large prompts, improving TTFT consistency for all users.