Now live: pip install galahad-kv
Galahad is the memory layer for AI. A model reads a text once, and Galahad keeps that reading. When the text is needed again, Galahad gives the reading back, so the GPU never reads the same text twice. Galahad also keeps the documents themselves and finds the part a question is about, so the model reads only what matters.
Galahad has a C++ core and runs inside llama.cpp, vLLM and SGLang.
- Don't pay for the same tokens twice. The tokens a model already processed are not processed again. Galahad gives the reading back instead of recomputing it. 99.6% of tokens came back from memory; 14× faster on vLLM, 0.59 s per question.
- Byte-exact, not approximate. What comes back is identical to what went in: the model's exact reading, and your documents stored as exact text. No lossy embedding, no "close enough" like a RAG pipeline. 100/100 right vs 77 for RAGFlow.
- Reads only what matters. Blaise finds the chapter a question is about, so the model reads 670 tokens instead of 9,700.
- Read past the context window. A text far larger than the model's context, up to 50 million tokens measured, is read in parts; each part's reading is saved, and the part a question needs is given back. GPU memory stays flat at 34.1 GB whether the text is 1 million or 50 million tokens.
- Memory that survives restarts, encrypted at rest, and your key never leaves your machine.
Free for 1 GPU.
pip install galahad-kvThen activate it (one command for everything):
galahad free --org <your-secret-key> --email you@company.com --accept-non-commercial
galahad doctor # library loads? licence valid? store writable?- vLLM: vllm serve <model> --kv-transfer-config '{"kv_connector":"GalahadConnector","kv_role":"kv_both"}'
- SGLang: add --hicache-storage-backend dynamicwith the Galahad backend (see the wiki)
- Bare / your own code: link libgalahad.so(ships in the wheel)
Linux x86-64, Python 3.10–3.14. On PyPI. Full guide: the wiki.
How it works
- Taliesin saves the model's KV cache to disk and restores it. It plugs in as the vLLM KV connector, the SGLang HiCache storage backend, and llama.cpp slot save and restore.
- Blaise keeps your documents as exact text and returns the chapter a question is about. Text is stored byte-exact.
- 50M-token window: a long text is read in parts; each part's reading is saved to disk and the part is dropped from the GPU, so GPU memory stays flat while the window moves over the whole text. The part a question needs is given back and answered from. Nothing to switch on. See the wiki.
- Inside your own program: llama.cpp is built into the Galahad library; your program calls its C API directly.
How it works
- Agent connectors: one decorator per step, for LangGraph, LangChain, CrewAI or your own loop. Each run becomes a causal graph.
- Replay: a pytest plugin. Record once; every CI run compares with the recording, so a prompt change that breaks the agent fails the build.
- Fault finder: returns the layer, the rule that decided it and the suspect block. With too little evidence it says "unknown" instead of guessing.
- Branching: copy-on-write branches of the KV cache (fork, snapshot, restore, discard). On vLLM and SGLang through a per-request galahad_branchflag; snapshots survive a restart.
- Inspector and rewind: a local web page and API behind a token, with an audit log. Rewind must be switched on by the host.
How it works
- Prefix sharing: the shared start is saved as a chain of chunks, and a lookup finds the longest saved start. It is used only when loading is faster than computing.
- Pinning: pinned blocks are skipped when Galahad has to free space.
- Tenants: each tenant has its own keyspace, with hits, misses and evictions counted per tenant.
How it works
- Encryption at rest: AES-256-GCM on every saved block and on the catalogue, with keys from your own key service. With a wrong key, the store refuses to open.
- Sanitizer: checks for NaN and infinity in the number format the host declares (bf16, f16), and checks the size.
- FinOps: uses your GPU price per hour and usage level. Without a price it gives no number and says what is missing.
Measured with Gemma 4 31B on llama.cpp, vLLM and SGLang, September 2026.
Galahad is tested on 30 open models from 10 families, from 7B to 70B.
© 2026 Corbenic AI - Sietse Schelpe. Patent pending.