Chutes released a year of production LLM serving request traces covering 6.12 billion requests across 9,174 models, enabling research into batching, scheduling, and GPU optimization. Key findings show high temporal locality in user requests, LRU cache effectiveness, and cache-aware routing improvements.