Lossless AI compression: 33% less GPU memory, bit for bit.
Glyd stores the weights of open models like Llama, Qwen, Gemma and Mistral in about 11 bits instead of 16, and decodes them inside the GPU's matrix multiply. Every weight comes back exactly, so you run the same model on less hardware, and often faster.
Glyd stores the weights of open models in about 11 bits instead of 16 and decodes them inside the GPU's matrix multiply. The same model on less hardware, often faster.
Measured on open models from
- Qwen
- Meta
- Mistral AI
- DeepSeek
- Hugging Face
- Microsoft
- Z.ai
Will it fit on your GPU?
Pick your GPU's memory. Weights plus an 8K-token KV cache and 1.5 GB for the runtime, bf16 against Glyd.
Weights, an 8K-token KV cache and the runtime, bf16 against Glyd.
Smaller, and usually faster
Generating a token reads every weight once, so reading a third fewer bytes saves time. At many sequences a step on A100 and H100 the decode work shows, and Glyd is slower there today.
All benchmarksDecode steps only, q, k, v and gate, up merged as vLLM runs them. Sep 26, 2026.
- 01Find the wasteA bf16 weight spends 8 bits on its exponent, which carries about 2.6 bits of information in a trained model.A bf16 weight spends 8 bits on its exponent, which carries about 2.6 bits of information.
- 02Code it exactlyCommon exponents get short codes and rare ones a side list. Sign and mantissa stay as they are: about 11 bits a weight, nothing rounded.Short codes for common exponents, a side list for rare ones. Nothing rounded.
- 03Decode in the multiplyThe GPU reads the compressed bytes and decodes them in registers, straight into the tensor cores. No bf16 copy is ever made.Decoded in registers, straight into the tensor cores. No bf16 copy in memory.
How it compares
Lossless weight compression is an active research area. Here is where Glyd sits, on numbers we measured ourselves where we could.
- Report · Sep 27, 2026Nineteen open models: every matrix 32 to 33% smaller, bit for bitEvery Linear layer's matrix of nineteen open models, SmolLM3 3B to GLM-4.5-Air, packed and unpacked bit for bit: 31.9 to 33.0% smaller, quality on ten.
- Report · Sep 27, 2026Qwen3.8 27B, the top open model for one GPU, in 41 GBQwen3.8 27B, the top-scoring open model that fits one GPU, in 41,071 MiB of GPU memory with Glyd against bf16's 51,771: under a 48 GB card, bit for bit.
- Plans · Sep 27, 2026What's next for Glyd: pip install, vLLM, one commandFrom a PyTorch harness to pip install "glyd[gpu]" with models ready on Hugging Face, then serving through vLLM, then glyd fit, pack and serve. No dates.
Is this quantization?
No. Quantization rounds weights to fewer levels, which changes the model. Glyd stores the same values in fewer bits and gives every one of them back. It is closer to a zip file than to 4-bit.
Does it change the model's answers?
The weights are the model's to the bit. The matrix products add their terms in another order than cuBLAS does, as any two GPU kernels do, so a close call can go either way: over the models measured, Glyd's MMLU answers are bf16's on 98.0 to 100% of the questions and perplexity stays within 0.06% of bf16's.
Which GPUs?
Measured on an RTX 4080 SUPER, RTX A6000s, an A10, an A100 and H100s. Glyd picks the layout for the GPU: the tiered one on RTX 40-series (Ada) and wherever only it fits, the 12-bit one on the others. Blackwell GPUs are not measured yet.
How do I install it?
Today Glyd on the GPU runs from its PyTorch harness: clone the repository, run gpu/setup_env.sh, then e2e.py on a model. A Python package, pip install "glyd[gpu]" with glyd.from_pretrained(), is in progress. The codec underneath installs with Homebrew, cargo or a wheel from the releases.
Try it on your model
Point the harness at any model released in bf16. It packs the weights, checks every one bit for bit, and runs bf16 and Glyd side by side on your GPU.
git clone https://github.com/surya-koritala/Glyd
cd Glyd/gpu
python e2e.py /models/Llama-3.1-8B-Instruct \
--format auto --fused --baseline --batch 1,8,32