Datasets:

Parameters Size

Typed Decisions

A benchmark for typed probabilistic decisions. A model gets one piece of unstructured state and answers five typed questions about it at once, and every answer is a probability distribution, not a single label.

The schema follows the System One primitives (noul, choice, score) used by

TypeSafe AI, so a row replays against any API with that

shape. The benchmark is independent: it is not affiliated with TypeSafe and does

not reproduce their Jev model.

Leaderboard

test split: 400 cases, 2,000 decisions. This table lists general models scored

zero-shot: they have never seen these workflows or their question schemas.

Models that were fitted or fine-tuned on train are in the second table below;

the two tables are not comparable.

Every row was scored by sending the whole case (state plus all five questions) in one request. Request shape matters: in a third-party run Jev's yes/no accuracy was 0.843 with one question per request and 0.788 alongside the others.

Hosted rows are p50 end to end from a client, one request at a time. † Measured on the same machine as the model (M3 Max), not comparable with hosted latency. ‡ Reported by the submitter on their own hardware (a different GPU and a different unit for each), not comparable with the other rows. § Self-reported by the submitter in the linked discussion and not re-run by us. Metric definitions follow the submitter's report; ECE in particular is computed differently by different submitters.

Fitted or fine-tuned on train

These models were trained on this benchmark's train split, so they are not

zero-shot and their scores are not comparable with the table above. All were

self-reported in the linked discussions and were not re-run by us, except the two

specialists marked ours and the three Bekko rows, which we scored ourselves. Scores well above the 0.735 teacher self-agreement mean a

model is learning the teacher's quirks.

¶ The submitter's own ECE definition, which does not match the one used above.

meraGPT Decider 1 is state of

the art among zero-shot models on this benchmark. It leads Jev on every question type (noul 0.840 vs

0.775, choice 0.733 vs 0.720, score 0.739 vs 0.696), its distributions sit far

closer to the gold (KL 0.096 vs 1.442), it is faster end to end, and it costs less

per token. It answers POST /v1/systemone at

meragpt.com, so the typesafe-sdk works

against it by setting TYPESAFE_BASE_URL.

To add a model, score it on test with the full distributions and open a

discussion with the numbers and the mode (specialist or general) it used.

Notes on the rows

- Liquid AI d1 was measured on 2026-09-30 through Liquid's API

(https://api.liquid.ai/v1/systemone,model: d1:free), with the same client code as the Jev row: all 2,000 decisions, zero errors. It ties Decider 1 onnoul(0.840) andchoice(0.732 vs 0.733) and trails onscore(0.677 vs 0.739), and it beats Jev on accuracy and KL. Liquid lists no per-token price yet, so the row shows the free tier.

- Jev 1.13.0 was measured on 2026-09-18 through TypeSafe's API (jev-latest, which reportedjev-1.13.0): all 2,000 decisions, zero errors, $0.016 in total. Its accuracy is near the 0.735 ceiling, but it puts nearly all its probability on one answer, which is where the KL gap comes from. Its confidence is not badly calibrated (overconfidence +0.023); the gold is a three-sample spread it does not reproduce.

- Jeff (firelex/jeff, Apache-2.0) was scored

on 2026-09-29 through the same client code as Jev, against Jeff's own

jeff-serve(commit2c1bfce) with the calibration each checkpoint ships. All three clear the Prior on accuracy but not on KL. Jeff's own 83.1% comes from a different five-benchmark panel. If there is a better way to serve them, open a discussion and we will rescore.

- Self-reported rows (§ above, and the second table) come from discussions #5 (prima-ratio), #7 (Bongard-mini), #8 (OpenDecider), #4 (od1), #3 (soft-decider) and #2 (Laya). Each submitter states the mode; we have not re-run them. If a number looks wrong, say so in the discussion and we will correct it.

- Bekko System One v0 (hotchpotch) was scored by us on 2026-10-01 on CPU (M3 Max, one call per case with all five questions, pinned revisions b886a1f917M,ab7685f268M,1960df56400M), with the same scorer as the other rows. Its training data includes this benchmark'strainsplit, with our teacher labels, so it sits in the second table; none of our 400 test cases are in its training data. We mapped our option and rubric text onto its candidate format in our option order. Its model cards say the license is not yet finalized. The author invited a score on X; we are glad to re-score with a different input mapping if he prefers.

- Specialists use Adaptive Classifier

0.2.0, one classifier per question on a frozen encoder, tuned on a held-out

quarter of train(mean pooling,max_length512, 30 epochs,prototype_weight0.3). Each case is entered four times, split across labels in proportion to its gold, so the soft target survives hard-label training; that cut KL by a third.

- Prior answers each question's trainlabel frequencies for every case. It has the best ECE while knowing nothing, which is why KL and Brier are the columns to read, not ECE.

Reading a score

Gold is the mean of three samples from a teacher of roughly 4B-class capability, so a score measures agreement with that teacher, not correctness. A better model can score lower wherever the teacher is wrong (it missed a duplicate invoice whose ID matched an earlier one).

Scores well above 0.735 mean a model is learning the teacher's quirks. Per-question

ceilings vary from 0.560 (agent_trace/urgency) to 0.937

(customer_service/category), so read scores per question as well as on average.

Specialist and general numbers are not comparable. A specialist is fitted on

train for these four workflows and cannot answer anything else. A general model

takes any question schema at request time and has never seen these. Say which

mode you used; the gap between them is the price of generality, not a quality

ranking.

The data

Every option carries a written description in criteria, and the descriptions

are part of the input.

state and questions together are exactly the body of a POST /v1/systemone

request. gold holds the full gold distributions, and flat

<question>__label / __probabilities / __score / __probability_true columns

hold the same answers for convenience. factors and label_agreement describe how

the case was built and are not model input.

from datasets import load_dataset

import json

ds = load_dataset("LocalLLaMA/typed-decisions", "customer_service", split="test")

row = ds[0]

state, questions, gold = (json.loads(row[k]) for k in ("state", "questions", "gold"))

print(gold["urgency"]["probabilities"]) # score against the full distribution

Report KL or log loss and Brier next to accuracy; calibration is the point.

How it was built

Each case starts from independently sampled latent factors (topic, tone,

severity, discrepancy type and so on), which a model renders into free text where

the artefact is textual and keeps structured where it is structured. A teacher

labels each case three times at temperature 0.7, and the gold is the mean of those

distributions, so it stays soft where a decision is genuinely ambiguous. Before

release, an audit checks state diversity, label balance, and that the gold

actually tracks the input. train comes from a separate run at a different seed;

packaging refuses to build if any case id or state appears in both splits.

The write-up behind the benchmark, including the two bugs it exposed in the classifier library: Typed Decisions on Latent Node.

- Downloads last month

- 21,940