SWE-Serve is a benchmark that evaluates agentic software engineering by converting real SGLang inference-serving work into coding tasks. Agents receive task instructions and a pinned repository, then must implement fixes using only local tools and Hugging Face access, with performance measured by pass rates on hidden functional and regression tests.