Quail is an execution engine for AI-SQL that extends SQL with LLM function calls. The article explains how to estimate latency of AI-powered filter operations using the roofline model, which computes arithmetic and memory requirements then divides by GPU hardware limits to obtain speed-of-light estimates for query execution.
This article discusses optimizing SQL queries that invoke large language models (LLMs) for data processing. The authors propose jointly optimizing query plans and LLM inference to achieve up to 14x speedups, addressing the high cost of AI-SQL systems that can generate millions of model calls per query.