On September 15, 2026, TypeSafe AI launched Jev and introduced a new category of models called System One models. Instead of generating a response, these models take a state, such as a support ticket, together with a set of possible options, and return a probability for each option in a single forward pass. There is no token generation and no output to parse.
At first, this looks very similar to traditional ML classifiers, but there is an important difference. In a conventional classifier, the labels are part of the model and the model is trained for that specific classification task. System One models receive the options as input, so the same model can be used with different sets of decisions without retraining.
Jev is closed, but open-source implementations appeared remarkably quickly. Within roughly two weeks, several alternatives were already available, including Laya, Lev and CLM.
Despite the different interface, these models are still built on pretrained language models: Lev and CLM use Qwen models as their backbone, while Laya is built on ModernBERT. The absence of token generation changes where the answer is read from the network, not the network itself. The underlying architecture still determines how the model processes the input and what behavior can be expected from it.
This makes architecture an important part of model evaluation. Benchmark accuracy shows how a model performed on a specific dataset, but not how its design affects its behavior across different tasks. This analysis derives testable hypotheses from the architecture and source code of three open-source implementations, then evaluates them through controlled experiments and statistical analysis. The same tests are finally applied to Jev, to infer the architecture of a model whose weights are closed.
The Models Architecture
All three models take the same two inputs: a state, which contains the information to be classified, and a set of options, which are the possible decisions. Each option has a name and, optionally, a description of what it means, such as “billing: charges, invoices and refunds”. The models return one probability for each option. Internally, however, the three implementations take very different paths to produce those probabilities.
Laya (Convai Innovations, 421M ModernBERT)
Laya first concatenates the options and the state into a single sequence, always in the following order: options first, state last. Each option is preceded by a [MASK] token, which works like the markers (A, B, C…) in a multiple-choice question and is used by the model to identify the options.
The model then processes this sentence as a standard ModernBERT would: it splits the text into tokens, converts each token into an embedding, and passes all of them through its 28 layers in a single pass.
Because the options and the state are processed together, the embedding of each option also incorporates information from the state during the pass. By the end, it reflects how well that option matches the state. This is where Laya departs from a usual LLM, which would use the final embeddings to generate the tokens of an answer. Laya stops at this point and keeps only the embeddings of the options. A small MLP converts each of them into a score, and a softmax turns the scores into probabilities.
Lev (Interfaze AI, Qwen3.5-4B + LoRA)
Lev is built on Qwen3.5-4B, a decoder LLM, and turns the task into a multiple-choice question. Like Laya, it concatenates the state and the options into a single sentence, but in the opposite order: state first, options last, each option labeled with a letter (A, B, C…). The sentence ends with an answer cue, such as “Answer:”.
The model then reads this sentence in a single pass, as Qwen normally would. Unlike Laya, Lev does not stop at the embeddings: it goes one step further and computes, for every token in its vocabulary, the probability of that token coming next. This is the step an LLM takes right before writing each word of an answer. Lev keeps only the probabilities of the option letters and rescales them so they sum to one. These become the option probabilities, and nothing is generated. To adapt Qwen to this format, Lev adds a small LoRA adapter to its layers, trained to push the probability toward the letter of the correct option.
To reduce the effect of the option’s letter on the prediction, Lev repeats the process with the options in a different order and averages the two results.
Lev also includes a second mechanism. I refer to the default mechanism described above as mode A, and to this one as mode B. In this mode, the options are not placed in the sentence. Instead, each one is converted into an embedding separately, which can be computed once and reused, and then compared with the embedding of the state.
CLM (Contrastive-LM, Qwen3-8B + two MLPs)
CLM is built on Qwen3-8B and, unlike Laya and Lev, never puts the state and the options in the same sentence. It first processes the state on its own, as Qwen normally would: it splits the state into tokens, converts each token into an embedding, and passes all of them through Qwen’s layers in a single pass. It then repeats the same process for each option, separately. When an option has a description, CLM encodes only the description; the name is used only to label the answer. With five options, for example, CLM runs six independent passes: one for the state and one for each option. This is the same principle as Lev’s mode B, used here as the only mechanism.
At the end of each pass, CLM keeps a single embedding that summarizes the whole text. These embeddings capture semantic similarity: two texts about the same subject are close to each other. That is not enough to make the decision. A state about a duplicate charge, for example, is close to several payment-related options, but only one of them is correct.
CLM therefore passes the embeddings through two small MLPs, one for the state and one for the options, which transform them into a new space of 512 numbers. The MLPs are trained on states paired with their correct options. As a result, closeness in this new space no longer means “same subject” but “correct option for this state”. CLM measures this closeness with the cosine similarity between the state and each option, and a softmax converts the scores into probabilities.
While Lev adds a LoRA adapter inside Qwen’s layers, CLM leaves Qwen untouched: the two MLPs, applied after Qwen, are the only trained part. And because each option is processed on its own, its embedding does not depend on the state. It can be computed once and reused for every new state.
The experiments below are built to leverage the architectural differences between the models.
The Experiments
To keep the comparison rigorous and reproducible, all three models were run under the same conditions: NVIDIA T4 GPUs in float16 precision, with the same states from the banking77 test split, the same option sets and the same random seeds. Laya and Lev ran on a single GPU; CLM’s larger Qwen3-8B was split across two. Every comparison is therefore paired. Latency values are the median of 30 measured calls, after warm-up calls that are not counted, with the 90th percentile recorded to show the spread. Before each experiment, the expected behavior of each model was written down from its architecture and source code, and each model’s output was compared with the example in its own README (see Notes for CLM).
One caveat applies throughout. According to its model card, Lev was trained on banking77, so its absolute accuracy is inflated. The comparisons below therefore focus on how each model’s behavior changes between conditions, which the training data does not explain. The minimal-edit experiment uses hand-written sentences and is not affected.
Experiment 1: Scaling With the Number of Options
The first experiment tests how latency and accuracy change as the number of options grows. Laya and Lev’s mode A write every option into the input, so each additional option makes the sequence longer. CLM and Lev’s mode B encode each option separately, allowing the embeddings to be cached and reused. The prediction is therefore straightforward: latency should grow with the number of options for the first group, but remain approximately constant when cached embeddings are reused.
Each of 100 states from the banking77 test split was paired with its correct intent and N − 1 distractors, for N = 2, 10, 50, 150 and 284. Intent names were used without descriptions, and the same option sets were given to every model.
For CLM and Lev mode B, latency was measured both cold, with an empty cache, and warm, with option embeddings already cached. Laya was run in two configurations: default, with its shipped budget of 512 tokens, which cuts the options when they do not fit, and fit, with the budget raised so that every option fits whole.
Absolute latency is not directly comparable across models because they use different backbones and hardware. Lev also ran without its optimized linear-attention kernels on the T4. The relevant comparison is therefore the shape of each curve.
The results follow the architectural prediction. With a warm cache, CLM and Lev’s mode B remain almost flat (mode B was tested up to 150 options; at 284 it ran out of GPU memory). Without the cache, both grew by about 10× because every option had to be encoded again.
Lev mode A grew 18.6× from 2 to 284 options, following the increase in tokens processed. Laya’s fit configuration grew 2.8×. Its default configuration appeared flat, but for a different reason: it truncates the options to keep the input within its 512-token budget.
At 150 options, Laya’s default configuration rejected 99 of the 100 questions. The one it accepted had room for the options but none for the state: all 32 state tokens were removed, yet the model still returned a probability distribution without warning.
The scaling result therefore has two distinct behaviors. Caching removes the cost of repeatedly encoding the options, while truncation can hide the cost by removing information from the input.
Accuracy reveals a second effect. Laya’s default and fit configurations are identical up to 10 options, but at 50 options truncation reduces accuracy from 0.63 to 0.52. Even without truncation, accuracy already falls to 0.63 at 50 options, while the sequence is still within the 512 tokens Laya was trained on, which points to a second mechanism; Experiment 2 examines it.
Lev mode A remains the most accurate, declining from 1.00 to 0.85, although its absolute accuracy is affected by its banking77 training overlap. Mode B starts at 0.69 with two options, consistent with a head trained primarily on larger label sets.
CLM falls from 0.64 with two options to 0.20 with ten and close to zero from 50 onward. The implementation was validated against the official vLLM path (see Notes), so this is a property of the model on this task. CLM was trained on answers to questions and on actions from agent trajectories, which are long texts, while the options here are intent names of two or three words. Their embeddings end up too similar to each other for the correct one to stand out, and the problem grows with the number of options. Experiment 5 shows this directly in the embedding space.
Experiment 2: Sensitivity to the Order of the Options
The options given to a System One model have no natural order: a support team listed third is the same team as when it is listed first. A well-behaved model should therefore return the same answer however the options are arranged. Whether it does depends on how the options enter the model. CLM and Lev’s mode B encode each option separately, so the order of the list should not affect them at all. Laya and Lev’s mode A read all options together in one sequence, where each option occupies a different place depending on the order, so the order can affect their decisions.
For each of 100 states from the banking77 test split, the model answered the same question five times, with the same options in five different orders: the original one and four random shuffles. This was repeated with 4, 10, 20, 50 and 77 options; with 77, the list contains every banking77 intent. As described earlier, Lev normally reads the options in two orders and averages the results, so it was also run with this step turned off, to see how much of its stability comes from it.
Both measures used in this experiment compare what the model returns for the same state when the options are shuffled. Suppose that, for one state with three options, the model returns these probabilities in two different orders:
The model chose A in both orders, so its answer did not change. The flip rate counts exactly this kind of change: it is the share of states for which the answer was not the same in all five orders. A flip rate of 20% means that, for 20 of the 100 states, shuffling the options changed the answer at least once.
The probabilities, however, did change: A fell from 0.8 to 0.6 and B rose from 0.1 to 0.3. The probability shift measures this. It is half the sum of the absolute differences between the probabilities of each option in the two orders (a measure known as the total variation distance in statistics):
where and are the probabilities of option in each order. In the example, the probability shift is (|0.8 − 0.6| + |0.1 − 0.3| + |0.1 − 0.1|) / 2 = 0.2: 20% of the probability moved from one option to another. The value is 0 when the two distributions are identical and 1 when they have nothing in common. For each state, it was computed for every pair of the five orders and averaged.
Shuffling the options
CLM never changed its answer, and its probability shift is exactly zero: each option embedding is computed once and reused in every order, so the model receives the same numbers whatever the arrangement. Lev’s mode B behaves the same way; its residual shift, below one millionth, comes from floating-point rounding when the options are encoded in batches of different composition.
Lev’s mode A reads the options in sequence, but it was only slightly affected. Even without averaging, shuffling the options changed its answer for only 6% of the states, and turning the averaging on reduced this to 5%, so most of its stability comes from the model itself. Its high confidence contributes: Lev assigns high probabilities to its answers, partly because it was trained on banking77, and a model that rarely hesitates between two options has few close calls that a new order can tip.
Laya was strongly affected. With 77 options, shuffling the list changed the answer for 79% of the states, and almost half of the probability moved between options. The effect is not caused by the truncation of its default configuration: the fit configuration, which keeps every option whole, still changed the answer for 74% of the states.
The order also changed how often Laya was right. With 77 options, it answered correctly 32% of the time when the correct option was in the first quarter of the list and 62% when it was in the last quarter. The same pattern appears with 50 options and is absent up to 20. Lev and CLM show no consistent trend.
The order of the options, therefore, affects only the models that read them in sequence, and strongly only Laya, whose accuracy depends on where the correct option appears in the list.
Distance to the state
In Laya, the options come before the state, so the options at the end of the list are also the ones closest to it. The second part of the experiment measures which of the two matters: the place of the correct option in the list, or its distance in tokens from the state.
To separate them, the options were moved farther from the state without changing their place in the list, by making every option longer without adding any meaning. In a second version of the list, the same words were appended to every option, so “card arrival” became “card arrival: banking intent”, and “lost card” became “lost card: banking intent”. Since every option receives the same words, they give no hint about which option is correct. But each option now takes about three more tokens, so the fifth option of the long list ends up farther from the state than the fifth option of the short list. Both lists had 50 options, stayed within the 512 tokens Laya was trained on, and were tested with 100 states, with the correct option placed at 11 different spots in the list, from the first to the last.
The results show that what matters is the distance. When the correct option was fifth in the list, Laya found it 55% of the time with the short list and only 33% with the long one, where the same fifth option was 134 tokens farther from the state. When the correct option was at the same distance from the state, the two lists gave the same accuracy, even though it sat at different places in each list:
A logistic regression confirms this by measuring the effect of both variables at the same time, estimating how much each one changes the chance of a correct answer while the other is held fixed:
For Laya, every 100 tokens between the correct option and the state multiply the odds of a correct answer by about 0.71, while the place in the list has no effect. The reason is how ModernBERT reads a sequence: in two out of every three layers, each token only sees the 128 tokens around it, and only every third layer sees the whole sequence. An option close to the state can read it in all 28 layers, while an option far from it can only read it in about 9. Accordingly, Laya’s accuracy stays at its highest, between 0.64 and 0.73, while the correct option is within about 45 tokens of the state, and declines between 45 and 75 tokens, around the 64 tokens that each side of the local window reaches. The decline is gradual rather than a sharp step, because the layers that see the whole sequence still connect distant options to the state.
Lev went through the same test, with averaging turned off so that each option appeared in only one place, and neither variable had an effect. In Lev, the state comes first and the answer is read at the end of the prompt, from a position that sees every token, so every option is equally close to the point where the decision is made. This is not only because Lev’s accuracy was already high: the probability it assigned to the correct option, which had room to vary, stayed between 0.85 and 0.93 at every distance. CLM was not tested, since it never places the options in a sequence.
The same mechanism explains part of Laya’s accuracy decline in Experiment 1: the more options in the list, the more of them sit far from the state.
Experiment 3: Minimal Edits
In most real inputs, the detail that decides the answer is small. A customer charged twice has a different problem from one charged once, and a customer who does not want to cancel a subscription needs the opposite of one who does. This experiment measures whether each model changes its answer when a single detail of the state changes, and keeps it when only the wording changes. Here again, how the options enter the model matters. Laya and Lev read the options together with the state, so each option is scored with access to the exact words of the text, including the one that changed. CLM compresses the state into a single embedding before comparing it with any option, and two sentences that differ by one word are mostly about the same subject, so the edit may barely change that embedding.
The experiment uses pairs of sentences that differ by one small edit. In most pairs, the edit changes which option is correct; in a smaller group, used as a control, the edit is a paraphrase that should leave the answer unchanged. The edits cover six categories:
Each question had three options: the two options of its category and a third option, “other”, for anything else. Each option also had a short description, such as “duplicate charge: the same purchase was billed more than once”.
To avoid relying on a single pair of texts per category, each pair was written as a template with a blank, such as “I was charged once / twice for my ___”. The blank was filled with three different words (coffee, taxi ride, hotel booking), turning one template into three pairs. Each category had four different templates, giving 12 pairs per category and 72 pairs in total, plus 15 control pairs built the same way from five templates. Every pair was run twice, with the options in two different orders, which gives 24 runs per category.
All texts were written for this experiment, so none of the models could have seen them during training. Laya ran in its fit configuration, so that no option was truncated. Since every option had a description, CLM encoded only the descriptions, as explained in the architecture section.
Two measures describe each pair. Take the first row: the correct answer is “single charge” for text A and “duplicate charge” for text B. A model that answers “duplicate charge” to both texts is right about text B, but only because it gives the same answer regardless of the text. For this reason, the main measure is the share of runs in which the model answered both texts correctly. The second measure shows whether the model noticed the edit even when its answer did not change. Suppose the probability of “duplicate charge” is 0.30 for text A and 0.45 for text B: the model registered the edit and moved in the right direction, but if another option kept a higher probability, the answer stayed the same. This difference, 0.15 in the example, is the shift toward text B’s correct option.
With 24 runs per category, each value has a margin of about ±10 percentage points, so small differences between categories should not be over-interpreted. The differences between models are large and consistent.
Lev followed almost every edit. Its probability moved strongly toward the correct option, by 0.59 to 0.95 on average, and it was the only model that kept every control answer unchanged. Lev reads the state and the options together, and the answer is read at the end of the prompt, from a position that sees every word of the text. Part of its advantage, however, does not come from the architecture: Lev is about ten times larger than Laya, and according to its model card it was trained on examples with swapped words and negated questions, which are close to the edits used here.
Laya also reads the state and the options together, and it registered the edits: its probability moved in the right direction in 79% to 100% of the runs. In many cases, however, the shift was not large enough to change the answer, and the individual templates show which edits Laya can follow. It succeeded when the edit introduced a word present in only one of the two texts, such as “twice”, “never” or “next week”. It failed when the edit changed a relation between words present in both:
In the first two templates, both texts contain exactly the same tokens in a different order, and Laya’s probabilities barely moved. The other three require comparing two quantities or tracking who did what to whom. In a single pass of a 0.4B encoder, Laya captures which words are present much more reliably than how they relate to each other. Lev solved most of these cases, but the flight template, which requires comparing two times, remained difficult for it too.
CLM failed almost every pair: at most 8% of the runs were correct on both texts in any category, and in 13 of the 24 templates it gave the same answer to both texts in every run. The reason is the separation between state and options. The edit changes only the state embedding, and in Qwen’s raw space it barely moves it: the probability of the correct option changes by only about 0.01, as expected for two sentences about the same subject. The MLPs amplify this difference to between 0.03 and 0.20, but in most cases not enough to change which option is closest. Since the option embeddings are fixed, any change in the decision has to come from this small change in the state embedding.
The controls show the opposite problem. Laya and CLM changed their answer in about one in five control pairs, so rewording without changing the meaning can move their decision. In Laya, the clearest cases were writing a number as digits or as words (“$200” against “200 dollars”) and replacing “sent” with “transferred”.
Following a small edit, therefore, requires reading the state and the options together. CLM, which compresses the state before seeing the options, lost almost every edit. Among the models that read them together, Laya followed edits that add or remove a word, but not those that change how the words relate, while Lev followed almost all of them, helped by its size and training data.
Experiment 4: What the Model Uses to Decide
As described in the architecture section, each option has a name and, optionally, a description. In normal use the two agree, so it is not possible to tell which one a model relies on. This experiment separates them. Each model’s architecture suggests a different answer. CLM encodes only the description whenever there is one, so the name never reaches the model. Laya reads each option at its [MASK], placed immediately before the name. Lev answers with a letter that points to a whole line, “A. name: description”, which the model reads entirely before giving its answer.
Twelve banking77 intents with clearly different meanings were used as options for 240 states, 20 per intent. The same states and options were presented in five conditions, changing only the text of the options. In some conditions, the name was replaced by a code, such as “K37”: a label with no meaning, which cannot tell the model anything about the option.
In the misleading condition, every option kept its own description but received the name of another intent.
Accuracy counts how often the model chose the option with the correct meaning: the option with the correct description or, when there is no description, the one with the correct name. With 12 options, chance is 0.083.
The misleading condition needs a second measure. Take the state “The ATM kept my card.” The correct option is the one with the description “a cash machine kept or retained the card”, which now carries the name “change pin”. Another option carries the name “card swallowed”, but the description of a different intent. A model that chooses this second option followed the name instead of the description.
A model can also choose that option by coincidence, simply because it made a mistake and that option happened to be the one picked. If a model is wrong in 55% of the states and its errors are spread evenly over the 11 wrong options, about 55% ÷ 11 = 5% of the states will land on the option with the correct name by chance. This coincidence level is the reference: a model really follows the name only if it lands on that option clearly more often.
With 240 states per condition, the accuracy values have a standard error of 2 to 3 percentage points.
CLM gave the same result, 0.4542, identical to the last digit, in the original, code and description, and misleading conditions. Since CLM encodes only the description, the model received exactly the same input in all three: only the name changed, and the name never reaches it. For the same reason, CLM followed the misleading name in 6% of the states, which is its coincidence level. With names only, it encodes the names instead and reaches about the same accuracy (0.44). Its low accuracy in every condition has the causes discussed in Experiment 1.
Laya followed the name. With misleading names, it chose the option whose name matched the correct intent in 72% of the states, eight times its coincidence level, with an average confidence of 0.80, so it was not hesitating between name and description. Laya can read the description: with a code in place of the name, the description alone kept an accuracy of 0.80. But when name and description conflict, the name prevails. This matches the architecture: the embedding that scores each option is taken at the [MASK], immediately before the name.
Lev followed the description. With misleading names, it kept an accuracy of 0.92, since it reads each option as a whole line before answering. The name still has a small effect: Lev followed it in 8% of the states, well above its coincidence level of 1%, because almost all of its errors fell on the option with the correct name. Here, however, architecture and training cannot be separated: Lev was trained on banking77 and knows these intents, and its training included shuffled options and paraphrased questions, which can by themselves teach a model to read the whole option.
With codes only, all three models fell close to chance (0.05 to 0.09): nothing in the state indicates which meaningless code corresponds to which intent. All three signaled this correctly, with an average confidence between 0.21 and 0.29, against 0.56 to 0.99 in the original condition.
The same set of options can therefore lead to different decisions depending on the model. For Laya, the names decide, so they have to be chosen carefully: a misleading or ambiguous name overrides a correct description. For CLM, the names are irrelevant as long as the descriptions are good. Lev reads the whole option and follows the description.
Experiment 5: Inside the Decision
The previous experiments observed the models from the outside, through their answers and probabilities. This one looks inside, at the point where each model reads its decision, to answer one question: what did the trained part of each model actually change?
That trained part is different in each model. In Laya, the whole ModernBERT was fine-tuned, and two extra layers and a scoring MLP were added on top. In Lev, a LoRA adapter was added inside Qwen’s layers, and it can be turned off to recover the original model. In CLM, Qwen was left untouched, and only the two MLPs were trained. For each model, the embeddings at the decision point were extracted with and without the trained part, for 300 states from the banking77 test split and the same 12 options as in Experiment 4, with names and descriptions. For Laya and Lev, these extracted embeddings gave the same answer as the models’ own prediction functions for every state, so they are the ones each model uses to decide.
Three measures describe what these embeddings contain. The AUROC answers whether the correct options can be told apart from the incorrect ones: 0.5 means not at all, and 1.0 means perfectly. The silhouette score answers whether the embeddings group by a given property, such as which option they belong to: positive when they group, close to zero when they do not, and negative when embeddings sharing that property end up apart. The probe accuracy answers whether the intent of the state can be read from an embedding by a simple classifier. The figures project the embeddings into two dimensions to show their organization via UMAP, while the numbers carry the evidence.
Laya
In the untrained ModernBERT, the embeddings group by option, and the correct options are spread across all groups: the model represents which option each embedding belongs to, but barely whether it fits the state. Its AUROC is above chance (0.67) only because each [MASK] can see the state through the bidirectional attention and pick up some overlap of words.
After training, the correct and incorrect options are almost perfectly separated (0.995), already at the output of the encoder. The two extra layers do not improve this separation; instead, they remove information that the decision no longer needs. The grouping by option turns increasingly negative, meaning that which option an embedding belongs to fades away, while the grouping by correct and incorrect tightens. What remains is essentially a single question, whether the option fits the state, which is what the scoring MLP reads. The figure still shows one strand per option in the encoder, because the projection preserves which embeddings are close neighbors, while the silhouette is dominated by the largest distances, which separate correct from incorrect options.
Lev
The LoRA adapter is not what makes Lev answer with letters. With the adapter turned off, Qwen3.5-4B already places 99.97% of its next-token probability on the option letters and chooses the correct one for 98% of the states. Reading the answer from letter probabilities is a capability the instruction-tuned model already has when the question is presented in its chat format, which is consistent with Lev’s model card, where the frozen model reaches 0.71 on its benchmark. Since the adapter was turned off, this result is not affected by Lev’s training on banking77.
What the adapter changes is the confidence and the internal organization. The probability of the correct letter rises from 0.965 to 0.998. Without the adapter, the intent of the state can only partly be read from the embedding at the answer position (probe accuracy 0.51, against a chance level of 0.08), and this embedding does not group by intent; as Figure above shows, it groups by the letter the model is about to say. With the adapter, the organization changes: the intent becomes almost perfectly readable (0.99), and the embeddings of states with the same intent group together. The adapter, therefore, shifts the embedding at the answer position from “which letter to say” to “which intent this is”.
CLM
In Qwen’s raw space, nearly all embeddings point in the same direction: any two states have an average cosine of 0.98, and any two options 0.95. This is a known property of using the last token of a language model as an embedding without training it for that purpose. In this space, the correct and incorrect options are practically at the same distance from the state (0.875 against 0.872), so choosing the closest option is close to chance.
The MLPs undo this concentration on the side of the states: the cosine between different states falls from 0.98 to 0.85, and states with the same intent group together, visible in the figure as single-color clusters that are absent in the raw space. On the side of the options, the effect is much smaller. The 12 options remain bundled together, away from the states, and more similar to each other (0.70) than to any state. The gap between the correct and the incorrect options widens only slightly (0.281 against 0.251), which is enough to raise the AUROC to 0.80, but the correct option is often not the closest one, so accuracy stays at 0.46. The options here are short descriptions of a few words, while CLM was trained on long answers and agent actions, and its option MLP does not spread such short texts apart.
The trained part, therefore, plays a different role in each model. In Laya, training is indispensable: the untrained encoder cannot tell a correct option from an incorrect one. In Lev, the backbone already answers correctly before any adaptation, and the adapter raises its confidence and organizes the embedding at the answer position around the intent of the state. In CLM, the two MLPs make the whole decision: they organize the states by meaning but leave short options too close to each other, which explains its low accuracy throughout this analysis.
Experiments on Jev
Jev introduced the System One category, and the three open-source models analyzed above were built in response to it. Its weights and code are not public, so it cannot be inspected the way the other three were. Most of the experiments, however, only require sending inputs and reading the probabilities that come back, and Jev can be queried through an API. Its architecture is unknown, but its behavior can be compared with that of the three models whose architecture is known.
Jev was queried through its API with the same texts, options and seeds as the open models. Before the experiments, it was checked against the example in TypeSafe’s documentation, a support ticket about shoes in the wrong size, and returned the expected option with probability 1.0.
Working through an API changed the method in a few places. The API accepts at most 255 options per question, so the largest set in Experiment 1 has 255 options instead of 284. Latency was measured at the client, so it includes the time the request takes to travel to the service and back; since this time is about the same for every call, it does not change the shape of the curves. Experiment 5 was not possible, since the internal representations are not accessible; the token counts reported for each call, described at the end of this section, take its place. Finally, Jev returns probabilities rounded to two decimals, so differences smaller than 0.01 between two answers cannot be observed.
Experiment 1: Scaling with the Number of Options
While the number of tokens grew by a factor of 7.4, latency grew by only 6%, from 247 to 262 ms. Jev’s latency is therefore dominated by a fixed cost of about 250 ms, which covers everything that does not depend on the input, such as the network, receiving the request, queuing and returning the answer. Reading about 2,000 additional tokens added only about 15 ms.
On the relative scale of Experiment 1, Jev’s latency grew by 1.06×, close to the flat curves of CLM and Lev’s mode B with a warm cache. Latency alone, however, cannot tell why. A model that caches the options and a model that reads them on fast hardware produce the same flat curve when the computation is small compared with the fixed cost. The next experiments tell the two apart.
Jev’s accuracy degrades slowly as the number of options grows, much less than Laya’s and CLM’s:
Only Lev is ahead, and Lev was trained on banking77. TypeSafe does not disclose Jev’s training data, so it is not known whether banking77 was part of it.
Experiment 2: Sensitivity to the Order of the Options
The open models always return the same probabilities for the same input, so any change between two orders could be attributed to the order. Jev offers no such guarantee. Each state was therefore also sent twice in its original order: the difference between these two identical calls measures the variation that has nothing to do with order, and the effect of order has to be read against it.
Two identical requests produced slightly different probabilities, and with 77 options the answer itself changed for 4% of the states. This is the usual behavior of a language model served on GPUs that process several requests together, where the order of the floating-point operations varies from one call to the next. The effect of order, however, is clearly larger: at every number of options, the probability shift between different orders is 2.5 to 4.3 times the shift between identical calls. Jev therefore does not encode each option on its own, as CLM does, since in that case changing the order would produce nothing beyond this variation. Its sensitivity is instead close to that of Lev’s mode A:
Jev’s probability shift is similar to Lev’s, while its flip rate is higher. This follows from Jev’s lower confidence on this dataset: with 77 options, its accuracy was 0.83, against 0.90 for Lev, and a less confident model has more close calls that a small shift can tip.
The place of the correct option in the list did not change Jev’s accuracy: with 77 options, it was 0.80, 0.84, 0.85 and 0.84 from the first to the last quarter of the list. The test with the short and long lists confirmed this. Without access to Jev’s tokenizer, the distance was measured as the number of characters of the options listed after the correct one. This span separates the correct option from the state in Laya’s layout and from the answer position in Lev’s, so it applies to either. Neither the distance nor the place in the list had an effect (p-values of 0.56 and 0.94), and the probability of the correct option stayed between 0.78 and 0.84 in every position of both lists. This is Lev’s pattern, and it rules out Laya’s layout, in which the options come before the state.
Experiment 3: Minimal Edits
Jev answered both texts correctly in every pair of five categories and in every control pair, where it never changed its answer. When an edit changed the correct option, the probability moved almost entirely toward it, by 0.96 to 1.00 on average. Jev also solved the flight template, which requires comparing two times and which Lev still failed.
Its only failure was the template “I entered the wrong PIN once”, which it assigned to “card blocked” in every run, as Laya and Lev also did. When three models that read the options together with the state make the same error on the same text, the text itself is the likely cause: the intended option, “pin question: a general question about the PIN”, is a poor fit for a sentence that does not ask anything. This template is better read as ambiguous than as a failure of the models.
Experiment 4: What the Model Uses to Decide
With the original options, Jev reached 1.00; with names only, 0.99; with codes and descriptions, 1.00; and with codes only, it fell to chance (0.06), with its confidence dropping to 0.13. The misleading condition, where name and description point to different intents, places it between Laya and Lev:
Jev followed the description in 56% of the states and the name in 43%, ten times its coincidence level, so both parts of the option carry real weight. Its confidence also dropped, from 0.99 with the original options to 0.71, so the conflict between name and description registered in its probabilities. This is the behavior of a model that reads the name and the description together, as a single line, rather than deciding mainly from one of them.
Experiment 6: What the Token Counts Reveal
Each response reports how many input tokens were charged for the call. These counts were identical in every repetition of the same request, so they measure exactly what TypeSafe counts. By changing one aspect of the request at a time, while keeping the others fixed, the counts show what Jev reads and how many times:
These results reveal four things about Jev.
- Jev wraps every request in a prompt of about 270 tokens. The minimal request, whose visible content has about 15 tokens, was charged 289, and extrapolating every series to zero content leaves the same fixed amount. A model that only encodes its input, as Laya and CLM do, needs no instructions, while an instruction-tuned language model such as Lev relies on them. This is the clearest indication that Jev is a decoder language model operating inside a system prompt.
- The options are written into the text the model reads, each with its own formatting. Each option costs about four tokens beyond its name, and descriptions are counted word by word. This is the pattern of a list inside the prompt in which each item carries a code or separator, like the “A. name” lines of Lev. The limit of 255 options, one less than 2⁸, is consistent with a fixed set of codes, one per option.
- The state is read once, however many questions the request contains. The long text had 81 more tokens than the short one, yet each additional question cost exactly 65 tokens with either text. If the state were repeated for every question, each question would cost 81 tokens more with the long text. TypeSafe’s documentation states that the questions are evaluated in parallel, and latency did not grow with the number of questions. Together, these point to a design in which the state is processed once and shared by all the questions, for example as a common beginning of the prompt followed by one branch per question.
- The output tokens are the size of the answer, not a computation. They grew by 8.7 tokens per option, which matches the size of each option–probability entry in the JSON returned by the API. They carry no information about the model.
These conclusions have one limit: the counts show what TypeSafe charges for, which is closely related to what the model processes but not necessarily identical to it.
Jev Inferred Architecture
Taken together, the observations describe a design close to Lev’s: a decoder that reads the state, then the options with a code for each, and derives the probabilities at the end. The other two designs are ruled out by independent evidence. CLM’s separate encoding of each option is incompatible with Jev’s sensitivity to order, and Laya’s layout, with the options before the state, is incompatible with the absence of any effect of distance. What Jev adds to Lev’s design, according to the token counts, is a state processed once and shared by several questions evaluated in parallel. Its accuracy and its handling of minimal edits, the best of the four models, are consistent with a larger or better-trained model than Lev, although its size cannot be determined from the outside.
This description is an inference from observable behavior. The weights, the prompt and the serving code remain closed, and a different architecture could in principle produce the same behavior. What the analysis shows is that the tests derived from the open models are enough to place a closed model among them, and that, on every test that could be applied, Jev behaves like the design in which a language model reads the options together with the state.
Conclusion
A benchmark score tells us how a model performed under one set of conditions. It does not tell us why it behaved that way, or whether the same behavior will hold when those conditions change. The experiments presented here show that these differences are not random. They follow from the way each model processes the state and the options, and from where its trained components sit in that process. That distinction matters beyond this particular comparison. Once the architecture is understood, its behavior becomes easier to anticipate. That can guide how a model is used, where its limitations need to be handled explicitly, and which parts of the system are worth changing when building a new application.
Laya is the right choice when speed and hardware matter more than precision. It is small enough to run on modest hardware and was the fastest of the open models. It is weak, however, in three situations, each caused by its design. With long option lists, its accuracy drops for options far from the state, because most of its layers only see nearby tokens. When the options do not fit its token budget, it cuts them and can silently drop the state. And because each option is scored next to its name, a misleading or vague name overrides a correct description. Laya works well for short option lists with clear names, while numeric comparisons are better done in code before calling it.
Lev is the most reliable of the open models. Because it reads the state and the options together and decides from a position that sees the whole prompt, it was barely affected by order or distance, it followed the descriptions, and it handled almost every minimal edit. Its weakness is cost: every question rereads every option, twice, so latency grows with the length of the list. Lev is the better choice when the decision depends on details and the option list is moderate; for large and stable lists, its mode B avoids the rereading. Its high accuracy on banking77, however, comes partly from training on that dataset, so it should be validated on the target domain.
CLM is good at what its separation of state and options guarantees: with cached options, its cost does not depend on how many options there are, and its answer never depends on their order. The same separation makes it weak at everything else tested here. The state is compressed before it meets any option, so small edits barely change the decision, and short labels end up too close to each other to be told apart. CLM is not suited to classifying text into short intent labels. It fits the task it was designed for: agents choosing among many stable actions, described in longer texts, where caching pays off.
Jev was the most accurate model and the most robust to minimal edits, and its design matches Lev’s with one addition: the state is read once and shared by several questions, so asking more questions of the same state costs little. Its weaknesses are those of a closed service: about 250 ms of fixed latency per call, small variations between identical calls, and no way to inspect or adapt the model.
These results also show that a System One model is not a new kind of network. Each of the three open-source is a pretrained language model, read at a point other than the generation of text: the embedding of a marker, the probability of a letter, or the distance between two embeddings. This makes the concept open to new variants, with even more innovative applications. Knowing why each design behaves as it does is what allows such variants to be built deliberately, rather than discovered by trial and error.
Notes
Reproducibility. The four notebooks, one for each open model and one for Jev, use the same texts, option sets and seeds, and the open models also share the same pool of 284 labels. Each notebook pins the exact commit of the model’s code. Repository with notebooks
CLM reader. CLM normally runs behind a vLLM server. In these experiments, it was replaced by a local reader that follows the repository’s recipe and plugs into the official CLM engine. Compared with the official vLLM path, it produced identical tokens, embeddings with a cosine similarity of at least 0.9999, and the same answers.
CLM’s README example. The README reports 0.9388 for “billing” on its support-ticket example. The official vLLM path, with the published MLPs, gives 0.9880, and the local reader gives 0.9874. Precision and serving stack were both ruled out, so the most likely explanation, not confirmed, is that the example was produced with an earlier version of the MLPs. This does not affect the experiments above.
Repositories.
- Laya (Convai Innovations): https://github.com/NandhaKishorM/laya
- Lev (Interfaze AI): https://github.com/InterfazeAI/lev
- CLM (Contrastive-LM): https://github.com/Contrastive-LM/CLM