A high-speed, local decision gateway running on Gemma 4 (llama.cpp / CUDA) that replaces flaky autoregressive JSON generation with direct token logit scoring, mathematical cyclic debiasing, asymmetric post-decision quotation, and Platt temperature calibration.
Get deterministic classifications, well-calibrated confidence scores, and verbatim audit trails in 15–45ms fast-path classification on consumer hardware.
gevva0-intro.mp4
Autoregressive LLM classification pipelines suffer from four critical failure modes:
- Logit Poisoning & Pre-Decisional Drift: Generating chain-of-thought (CoT) before categorical output allows early hallucinated tokens to corrupt the final decision probability.
- Positional Bias: Standard single-shot logit scoring shows steep label favoritism (e.g., models picking "A" ~70% of the time regardless of option content).
- Severe Overconfidence (Poor Calibration): Raw softmax distributions yield high confidence (95%+) on wrong answers, making automated routing thresholds risky.
- Latency Tax: Autoregressive JSON schemas require 100–300 generated tokens, adding 500–2,500ms of latency per request.
Standard Autoregressive Classification (500–2,500ms):
Prompt ──> Generated CoT / JSON ──> [Logit Poisoning Risk] ──> Brittle Output
Gevva0 Calibrated Dual-Path (15–45ms):
Prompt ──> Direct Logit Readout ──> Platt Scaled ──> Confidence >= Threshold?
│
├── YES ──> Fast-Path Locked Verdict (15-25ms)
└── NO ──> Bounded Verification (30-85ms) ─┘
│
[Verdict Permanently Locked]
│
▼ (Isolated Forward Pass)
Asymmetric Grounded Evidence Extraction
Evaluated across easy, original, and hard forensic legal batteries from the official JevBench v1.4.2 benchmark suite.
See full evaluation report:
docs/JEVBENCH_PUBLICATION_EVALUATION_26B.md
All local conditions were evaluated under strictly isolated weights (gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf, 14 GB) across 650 balanced multi-class scenarios (including 18.5% out-of-distribution distractor controls).
See full evaluation report: Gemma 4 26B-A4B (MoE):
BENCHMARK_REPORT_gemma-4-26B-A4B.md
Requires Python 3.12 and an NVIDIA GPU (CUDA 12+ / 13+). Prebuilt wheels are bundled via llama-cpp-python:
# Using uv (Recommended)
uv sync
# Or using standard pip
pip install -r requirements.txtVerify your GPU environment and model loading:
uv run gevva0 checkfrom gevva0 import GevvaEngine
# Initialize the engine (auto-discovers model from llm_config.json)
engine = GevvaEngine()
context = "Customer reports unauthorized double charge on order #89211 after payment timeout."
options = [
"Billing Dispute / Refund",
"Technical Bug / Gateway Timeout",
"Account Security Incident",
"General Inquiry"
]
# Run decision with cyclic debiasing and confidence calibration
result = engine.decide(
context=context,
options=options,
confidence_threshold=0.85,
cyclic_debias=True
)
print(f"Verdict: {result.selected_option}")
print(f"Confidence: {result.calibrated_p:.4f}")
print(f"Path Taken: {result.path}") # 'fast_path' or 'adaptive_cot'
# Asymmetric post-verdict quote extraction (zero logit poisoning)
audit = engine.extract_audit(context=context, locked_choice=result.selected_option)
print(f"Grounding Quote: '{audit.verbatim_quote}'")# Fast-path decision outputting structured JSON
uv run gevva0 decide `
--context "Customer reports their invoice was charged twice for the same month." `
--options "A: Billing Inquiry,B: Technical Bug,C: Churn Risk" `
--json
# Force CoT scratchpad via high confidence gate threshold
uv run gevva0 decide --context "..." --options "A: x,B: y" --threshold 0.99
# Full cyclic label-permutation debiasing (invariance guarantee)
uv run gevva0 decide --context "..." --options "A: x,B: y,C: z" --cyclic
# Launch API service & Web Dashboard UI (http://localhost:8000/ui)
uv run gevva0 serve --host 127.0.0.1 --port 8000Click to expand Architecture & Theoretical Foundations
-
Direct Token Logit Scoring & Prefix Marginalization: Evaluates candidate token log-probabilities directly from llm.scores[llm.n_tokens - 1]. Marginalizes over bare and space-prefixed tokens (logsumexp(logit(" A"), logit("A"))) to capture true prior distributions without running an autoregressive decoding loop.
-
Bounded System 2 Verification: If the fast-path calibrated confidence score is below the threshold, the gateway executes a micro-scratchpad (up to 40 tokens at $T=0.2$ ) focused purely on option exclusion, then immediately rescores the final choice.
-
Guaranteed Label Invariance (Cyclic Debiasing): Evaluates cyclic option permutations ($A \to B \to C \to D \to A$ ) and projects probability mass back to semantic labels, eliminating positional favoritism.
-
Platt-Style Temperature Calibration: Minimizes multi-class Brier score over validation scenarios to produce a learned temperature parameter ($T$ ), aligning raw logit softmaxes with true empirical accuracy ($ECE < 0.03$ ).
- Asymmetric Audit Trail (Post-Decision Extraction): Once the verdict is locked, an isolated, secondary forward pass extracts verbatim quotes and rationales from the source context. The rationale cannot retroactively corrupt or poison the classification decision.
For mathematical proofs, multi-class Brier score Murphy decompositions, and VRAM sizing charts, see
docs/ARCHITECTURE.md.
Click to expand Repository Layout & Configuration
├── src/
│ └── gevva0/ # Core Python package
│ ├── engine.py # llama.cpp logit extraction, KV caching, calibration
│ ├── debias.py # Cyclic permutation debiasing
│ ├── audit.py # Asymmetric quote & rationale extractor
│ ├── calibration.py # Platt temperature calibration & Brier fitting
│ ├── metrics.py # Bootstrap CIs, McNemar tests, ECE, Brier decomposition
│ ├── config.py # Model path, context size, and auto-discovery
│ ├── schema.py # Pydantic request / response schemas
│ ├── server.py # FastAPI service (serves /ui and API routes)
│ └── cli.py # CLI entrypoint ('gevva0')
├── ui/
│ └── web_dashboard/ # Interactive dashboard & audit UI (mounted at / and /ui/)
│ └── index.html # Single source of truth web interface
├── benchmarks/
│ ├── benchmark_suite_650.json # Rigorous N=650 4-class balanced evaluation battery
│ ├── benchmark_suite.json # Standard benchmark suite
│ ├── generate_rigorous_dataset.py# Generator for N=650 dataset with 18.5% OOD controls
│ ├── run_benchmark.py # Standardized CLI runner with 4 ablation lines
│ ├── generate_fixtures.py # Generates multimodal test images
│ └── fixtures/ # Test image assets (AP invoices, 404 UI, CCTV frames)
├── tests/
│ ├── tests.json # Curated interactive showcase tests for Web UI
│ ├── test_metrics.py # Unit tests for statistical & calibration formulas
│ └── test_benchmark_runner.py # Integration tests for benchmark runner
├── docs/
│ ├── ARCHITECTURE.md # In-depth technical write-up on logit scoring & calibration
│ ├── BENCHMARK_REPORT_gemma-4-26B-A4B.md
│ └── JEVBENCH_PUBLICATION_EVALUATION_26B.md
├── models/
│ └── gemma-4-26B-A4B-it-qat-UD-Q4_K_XL/
│ ├── gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf
│ └── mmproj-BF16.gguf
├── requirements.txt # Autogenerated requirements for pip / non-uv environments
├── pyproject.toml # Project dependencies managed via uv
└── README.md
{
"n_ctx": 4096,
"n_batch": 2048,
"n_seq_max": 8,
"kv_unified": true,
"kv_cache": "F16",
"type_k": 1,
"type_v": 1,
"model": {
"model_path": "models/gemma-4-26B-A4B-it-qat-UD-Q4_K_XL/gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf",
"mmproj_path": "models/gemma-4-26B-A4B-it-qat-UD-Q4_K_XL/mmproj-BF16.gguf",
"n_gpu_layers": -1,
"verbose": false
},
"decision": {
"cot_threshold": 0.85,
"cot_max_tokens": 256,
"cot_temp": 0.0,
"cot_prompt": "Analysis: First calculate everything: ",
"cyclic_debias": true,
"kv_branching": true,
"kv_branching_min_tokens": 50
}
}Overrides can also be set via environment variables: GEVVA0_MODEL, GEVVA0_N_CTX, and GEVVA0_CALIBRATION.
The benchmark suite includes an automated runner with
# Run benchmark on active model
uv run python benchmarks/run_benchmark.py
# Run on specific parameter scale
uv run python benchmarks/run_benchmark.py --model 26b # Gemma 4 26B-A4B MoE
uv run python benchmarks/run_benchmark.py --model e4b # Gemma 4 E4B Dense
uv run python benchmarks/run_benchmark.py --model e2b # Gemma 4 E2B Edge
# Re-generate synthetic N=650 balanced evaluation battery with OOD controls
uv run python benchmarks/generate_rigorous_dataset.pyApache License 2.0. See LICENSE for details.