BM25 is a probabilistic weighting scheme used by Xapian that combines previous models (BM11 and BM15) with a scaling factor to improve term weighting in information retrieval. Recent TREC tests have shown it to be the best known probabilistic weighting scheme, with default parameters of k1=1, k2=0, k3=1, and b=0.5 that can be tuned for specific collections and query types.
BreadBowl-Embed proposes a new retrieval representation that sits between single-vector and token-level approaches, using 16 fixed slots per passage with dual vectors for routing and value reading. This addresses the inefficiency of two-stage RAG pipelines where documents are read twice—first for retrieval via bi-encoder, then for reranking via cross-encoder—reducing computational waste especially for agents that issue multiple queries.
A comprehensive survey examining efficiency in large language model-based agents, focusing on three core components: memory, tool learning, and planning. The paper reviews approaches for reducing costs such as latency and tokens through techniques like context compression, retrieval, budgeted tool use, and hierarchical planning, while evaluating efficiency through Pareto frontier analysis between effectiveness and cost.
Zenith argues that AI agents need routing-based retrieval systems tailored to different question types rather than a single best retrieval method. Vector search alone fails for questions requiring document versioning, relationship paths, or current state; the solution is hybrid retrieval combining vector and BM25 search, with GraphRAG only when relationship queries genuinely outperform tuned baselines.
Researchers demonstrate a memory layer that extends AI language models to handle 50-million-token contexts by storing and retrieving key-value states from encrypted disk storage, achieving 2.8-4.3x faster loading and 8.8-12.3x lower GPU energy use compared to recomputation. Testing on Gemma models shows accurate retrieval of facts from millions of tokens earlier with no hallucinations.
Med Karim Bchini ports jeffhub.ai use cases to Google's EmbeddingGemma-2 embedding model, replacing a 0.8B decider with dense embeddings and cosine similarity scoring. The Node.js implementation runs CPU-only using quantized ONNX weights (~314 MB), achieving ~170 ms latency per embedding and 100% accuracy across classification and retrieval tasks on small curated datasets.