Typed Decisions is an open benchmark on Hugging Face for evaluating probabilistic decision models using a standardized schema with five typed questions per unstructured input. meraGPT Decider 1 leads zero-shot performance, outperforming TypeSafe's Jev on accuracy and KL divergence metrics across all question types.
Credence is a local inference runtime that makes typed probabilistic decisions using GGUF language models by scoring the model's next-token distribution over permitted labels without generating text. Built on llama.cpp and framework-agnostic, it returns boolean decisions with probability scores and uncertainty diagnostics, currently in Phase 1 with a 4B model achieving 22/22 accuracy after calibration.
TypeSafe AI released Jev, a fast probabilistic classifier for extracting structured answers from unstructured text, but it lacks true reasoning capabilities compared to symbolic reasoners like Rainbird. Jev works well as a boundary layer between messy language and formal decision systems, though developers must handle complex reasoning logic themselves.