Ternary-Bonsai-8B is a 1.58-bit quantized language model in GGUF Q2_0 format, using ternary weights ({-1, 0, +1}) with shared FP16 scales. The model ranks 2nd among compared 6B-9B parameter models despite being 1/8th their size, with implementation support in a custom llama.cpp fork.
This repository implements line-rate post-quantum Byzantine consensus using balanced ternary representation and L1D cache optimization, achieving 12.4 Mpps/core throughput at 100GbE with sub-80ns latency. The system comprises eBPF/XDP drivers for packet processing, Triton GPU kernels for efficient decompression, and microbenchmarking tools, targeting top-tier systems conferences.