A model can be wrong and unsure, or wrong and so confident that nobody checks.

A new kind of AI model, the decision model (also called a System One model), is built for

automated decisions. It doesn't write text. It picks from a fixed set of options and returns a probability for

each one. Software can act automatically when the model is sure and send the case to a person when it

isn't. We tested the best-known one, Jev from TypeSafe AI.

Picture a bank routing thousands of support messages an hour. Nobody reads each one, so the confidence number decides which ones a person sees. If the number drops when the model doesn't know the answer, the second model in Figure 1 gets stopped. If it doesn't, the mistake goes straight through. TypeSafe says Jev's confidence does drop. Its launch post promises:

"Calibrated decisions: answers with epistemically honest probabilities on System One

tasks."

TypeSafe, Introducing System One

models and Jev, accessed 29 September 20261

That sentence makes two promises. To see what they mean, here is one of our calls, as sent and as returned (token counts omitted):

// request

"state": {"question": "What number will come up on a single roll of a fair six-sided die?"},

"questions": {"answer": {"type": "choice",

"instructions": "Choose the correct answer to the question.",

"criteria": {"one": null, "two": null, "three": null,

"four": null, "five": null, "six": null}}}

// Jev's response

"model": "jev-1.13.0",

"answers": {"answer": {"type": "choice", "choice": "one", "confidence": 0.73,

"probabilities": {"five": 0.01, "six": 0.09, "two": 0.01,

"one": 0.77, "three": 0.06, "four": 0.06}}}

Jev picked "one" and gave it a probability of 0.77. The response also has a confidence score of 0.73, but that's not a separate judgement. It's the 0.77 rescaled so that spreading evenly over the six faces (1/6 each) would read 0 and certainty would read 1. The formula is roughly (6 × 0.77 − 1) / 5, which gives 0.73 up to rounding. Only the 0.77 can be checked against how often Jev is right, so in this post "confidence" means that number, the probability of the answer Jev picks.2

The first promise is calibration. Of all the answers Jev gives 77%, about 77% should be right. The second promise is about this answer alone. The die is fair, so each face has a one-in-six chance and nothing in the question favours "one". By saying 77%, Jev claims to know something it can't. This wasn't one unlucky call. We asked the same question with the six options in all 720 possible orders, and Jev picked "one" every time, with 0.80 on average. A model can keep the first promise and still break the second. If cases like this are rare, the overall numbers still look calibrated, yet each such case clears a gate set at 50%. The gate depends on the second promise.

We tested both claims in about 575,000 Jev calls.3 Four results stood out:

On familiar closed-choice tasks, confidence tracked accuracy well and supported useful filtering.

When the prompt contained no basis for an answer, or the question fell beyond the observed knowledge boundary, confidence often stayed high.

Asking "Do you know?" seemed at first to detect these gaps. Once we controlled for names that looked made up and for dates in the questions, it did no better than ordinary answer confidence.

Questions about the case itself held up better. Jev could tell whether the evidence it was given was enough, and whether an outcome was already settled.

Reading a probability on its own can't test the second promise. A 50% can mean the model is missing a fact, or that the outcome is a coin toss. The first is epistemic uncertainty, which more information can reduce. The second is aleatoric uncertainty, which no amount of information removes. So we built situations where we knew the answer. In each one we changed exactly one thing about what Jev was told and watched what its probabilities did. In other tests we stated the odds outright.

Our first test used customer-support messages from two public datasets, Banking77 and CLINC150. Take this one:

"I ordered a card and I still haven't received it. It's been two weeks. What can I do?" The right

category is "card arrival". We replaced every category name with a meaningless code, such as

INTENT_67, so that without help nobody could tell which code was which. Then we gave Jev

different amounts of help: nothing, a few examples per code, the real names, irrelevant text, or examples attached

to the wrong codes. Click through the settings:

With one example per category, Jev gets it wrong. The only example for "getting virtual card" was "my virtual card has not came yet!", which sounds a lot like our customer. Two examples fix it, and with eight Jev is 96% sure and right. Across all 2,270 messages, more information meant higher accuracy and higher confidence:4

Now click No hints. For our question this is the setting that matters most, because here the right answer is unknowable. Jev was still 32–36% sure of its top pick on average. An even split across the options would be about 1%, and that's roughly how often it was right. Mostly it went for whichever code looked like a default, such as INTENT_00:

We made this test artificial on purpose, so there is no doubt the answer was unknowable. It shows that Jev can sound sure with nothing to go on.

How the test was set up, and two comparison models

Jev followed swapped examples on 90–98% of items. So the swap changed the mapping Jev used. Scoring

those answers against the original labels is not, by itself, evidence of miscalibration.

We ran two comparison models on the same messages. An open model built for the

same job, GLiNER2.5-Decide, stayed close to an even split with no information (1.7% on its top option) but was

less accurate than Jev with category names. Five small models trained separately on the same examples (our

"students") were less accurate and less confident. With one banking example per category, their confidence ran 10

points ahead of their accuracy, against 15 for Jev.

The open model is fastino/GLiNER2.5-Decide, run locally without training. The

students are five fine-tuned copies of a 22M-parameter encoder that differ only in random seed. We use their

disagreement as a rough measure of missing knowledge, a common heuristic rather than ground truth.

2. Where confidence worked

None of this makes Jev's confidence useless. When Jev had what it needed, as on standard multiple-choice benchmarks, its confidence tracked its accuracy closely:

Take trivia questions (TriviaQA, turned into four-option multiple choice). Keep only the answers Jev was at least 90% sure of, and you keep 82% of the questions, with 99.75% of those answers right.5 That's a strong result for this dataset, though it's no guarantee for other tasks.

It wasn't perfect, even here. On Banking77, where some categories are near twins, Jev was 83% confident but only 68% accurate with one example per category. It was also overconfident on MMLU-CF, a benchmark built to avoid questions a model may have seen in training.

More charts: confidence against accuracy for intent routing, and quiz clues revealed one at a time

Questions about real people, films and places gave the same result. From the famous to the obscure, confidence matched accuracy. Then we asked the same kinds of questions about 1,600 made-up subjects ("Who wrote The Velbri Tide?"), where none of the four options is right. Jev had nothing to go on, and it still put about half its probability on one option.

3. Later news, higher confidence

Made-up subjects are still a bit artificial. News gives a natural test, because a model's knowledge stops somewhere in time. We asked Jev yes/no questions about real events from January 2020 to July 2026, like "Will X happen by March 2025?". Its accuracy dropped sharply around late 2024. We call that point the observed knowledge boundary. We found it in the data, and it isn't a confirmed training cutoff.

Before the boundary, Jev was right 70% of the time and 75% confident on average. Beyond it, Jev was right only 51% of the time, no better than a coin flip, yet its confidence rose to 82%. It answered "no" to 92% of the later questions, even though about half of those events did happen. It behaved as if not having heard of something meant it hadn't happened.

The dates in the questions drove much of that "no". When we took them out, Jev said "no" much less often to later events that did happen (57% instead of 92%), while earlier questions barely changed. Removing a date can change what a question means, though, so we read this as a clue to how Jev responds, not as a measure of accuracy.6

How we checked the date edits, and their limits

An AI assistant assessed 240 original and edited question pairs against a written protocol. These labels are

provisional and haven't yet been checked by a person.

About 35% kept their answer, 32% lost a deadline that could change a "no", and 33% no longer clearly identified

the event. On the pairs judged to keep their answer, later "no" answers fell from 92% to 65% (37 pairs).

Can recalibration fix it?

The usual fix is recalibration, which learns a mapping from the model's confidence to how often it's right. It works only if the same confidence means the same thing on new cases, and here it didn't. At the same confidence, later questions were right 17–31 points less often than earlier ones. A mapping trained on earlier months still left later questions about 20 points overconfident.7 The number alone carries no sign of the shift.

4. What we learned so far

Jev's confidence is useful on familiar ground and unreliable beyond it, and correcting the number afterwards

doesn't help. In Part 2 we ask Jev directly whether it knows, test what that question

reacts to, and look at the questions that did hold up.

5. Methods, limitations and provenance

Scope and measurement

One model version. Requests used the jev-latest alias; every response reported

jev-1.13.0. Other versions and model families may behave differently.

Scale. About 575,000 API calls (575,442 by the machine-generated ledger) across 15 public

datasets and six generated task families. A call can carry several questions, so calls, questions and unique

items are different counts.

Aggregation. Most items were sent three times; we averaged the returned distributions, then

took the top option and its probability. With one call per item instead, estimated calibration error changed

by less than about one point; that doesn't mean every individual decision would be the same.

Calibration. We report top-label calibration error (smooth ECE), confidence–accuracy gaps,

reliability charts and proper scores. Low top-label error doesn't imply calibration per class, per subgroup or

in deployment.

Intervals. Item-clustered bootstrap intervals, month clusters for dated news, and moving

blocks for analyses over time.

Multiple choice is easier than open questions. Most free-text benchmarks were converted to

four options with automatic distractors.

Construct validity and exploratory analyses

False premises. Made-up-subject questions have no correct option, so they have no

calibrated 25%-per-option target; an even split is only a forced-choice reference.

Observed knowledge boundary. The boundary is a statistical change in accuracy, not an identified

training cutoff; topics, style and difficulty can change over time.

Date edits. Removing dates can change a question or lose the event. The 240-pair assessment

was provisional, done by an AI assistant, with independent human checking pending.

Stand-ins for knowledge. Made-up/real status and period are not direct labels of what the

model knows.

Exploratory controls. The surface-cue controls, date edits, verification variants and

recalibration bounds were added after the main studies. The findings over time have not been confirmed on a

fresh set of news.

Decision value. The routing comparisons don't establish value for every mix of questions,

cost of errors or kind of shift.

First stage: key numbers for Jev

Accuracy, average confidence (top probability) and calibration error (smooth ECE: 0 means confidence matches

accuracy at every level; it is not the same as confidence minus accuracy) per setting. 770 Banking77 and 1,500

CLINC150 messages per setting, each the average of three identical requests. The swapped-examples row is scored

against the original labels; Jev followed the swapped examples, so it shows that the swap took effect rather

than a calibration failure.

Authorship and assistance. The authors are affiliated with Synthpop.AI. An AI

coding assistant helped implement the experiments and draft text, and produced the provisional date-edit labels.

The authors take responsibility for the study. Written critiques of earlier drafts were informal feedback, not

formal peer review.

Disclosure pending author confirmation: any commercial or financial

relationships with TypeSafe, TypeLLM or other relevant model providers.

Quoted from TypeSafe's release post, "Introducing System One models and

Jev" (accessed 29 September 2026), which also says Jev "always communicates confidence and uncertainty

with every output. Calibrated: higher confidence means higher accuracy." Its documentation describes Choice

probabilities as "the full probability distribution across every option" and System One probabilities as

"optimized against outcomes to reflect uncertainty" (docs.typesafe.ai, accessed 30 September 2026). "Honest"

here describes how the probabilities behave, not intent. ↩

We sent this request on 1 October 2026 (request ID req_01a0f72fbc7f7526b82d55876fd05fc3). The

open-source code rebuilds it exactly: it is unit device:die_words:0 of the

stated_odds experiment, in which all 720 orders gave the 0.80 average. In general the

confidence score is close to (K·p − 1)/(K − 1) for K options, so the same probability gives

different scores for different numbers of options: 0.77 reads about 0.72 with six options but 0.54 with two.

It carries the same information as the probabilities, and calibration needs the probability scale. ↩

575,442 paid API calls by our machine-generated ledger, about 1.15 billion input tokens on one model

version (jev-1.13.0), an estimated $48 at the listed input price. Public benchmarks were turned

into Jev's question types, usually four-option multiple choice with distractors from the same dataset. ↩

The spread of Jev's probabilities (entropy divided by its maximum) fell by 0.66 (Banking77) and 0.54

(CLINC150) from no examples to one, then by only about 0.02 and 0.01 for each further doubling. Three different

random draws of examples gave near-identical curves. ↩

8,132 of 9,960 questions kept, 20 wrong. A 90% threshold doesn't mean the kept answers are 90% right;

most sit well above it. Calibration error (smooth ECE, 0 is perfect) was 0.016 on MMLU-Redux, 0.031 on TriviaQA

and 0.010–0.026 for CLINC150 with 5 to 150 options. On MMLU-CF, confidence was 91% at 78% accuracy. ↩

The boundary is the single change point in monthly accuracy, at November 2024; the drop and the rising

confidence appear whichever month is used, and within every news category. The date test used 1,680 later and

1,680 earlier questions with every date phrase removed by rule ("Will the Fed cut rates by the end of March

2024?" became "Will the Fed cut rates?"). Telling Jev today's date, either the study date or a date just after

each event, changed neither later-event accuracy (51%) nor the "no" answers (91–92%). When asked to check

proposed answers, Jev favoured "it didn't happen" whether asked "is this answer correct?" or "is this answer

wrong?". ↩

Standard temperature and Platt scaling left the gap beyond the boundary at 24 points. Maps fitted on some

months and tested on others left 15–18 points; maps fitted only on earlier months left 19–23 points. An exact

bound shows no order-preserving map can match accuracy in both periods on this data. The recalibration appendix of the paper gives the full procedure. ↩