A new inference engine called Quail combines query planning with inference optimization to achieve over 1 billion tokens per minute on a single H100 GPU, delivering 1.84x faster performance than vLLM on AI-SQL queries. The system optimizes for structured data transformation workloads by intelligently managing key-value cache across requests, enabling cost-effective inference at under 6 cents per billion tokens.