ej 0.0.1 is a small calibrated decision model; technical report forthcoming.
Weights: https://huggingface.co/5ak3t/ej, revision
v0.0.1, one filemodel.ejpack(11,384,312 bytes, SHA-256e990e1846cba43f8405a969c606057f2fd6e2076595f4a34d202e8fc531891b0). Code, training and evaluation: https://github.com/5ak3t/ej (file paths below refer to that repository). Release notes:docs/releases/0.0.1.md.
A System-1 decision model for devices: given a state (text, or a JSON object as text) and typed questions, it returns
one probability distribution per question in one pass, without decoding tokens. Usage: Quickstart below and the
repository README.
- No LLM at inference; teachers are used at fit time only and their weights do not ship.
- Calibration is fitted, then checked per suite: whether the result is calibrated is measured on each suite against a perfect-calibration ECE floor (§5).
- Record independence. Zero-shot predictis record-independent: each record's output depends only on that record (up to the float noise of batching).adapt(andAdaptedModel.observe) is an opt-in, per-workflow transductive mode: it pools option statistics across that workflow's labelled records, so use it only for one workflow with a fixed option set (the same question ids and option keys).
- Option texts are read as text, so new label spaces can be asked zero-shot; quality on unseen workflows is limited (§6).
Architecture and method: technical report forthcoming.
git clone https://github.com/5ak3t/ej && cd ej
pip install -e .
import ej
model = ej.load('5ak3t/ej', revision='v0.0.1') # downloads model.ejpack only, checks its sha256, nothing else is fetched
(probs,) = model.predict([ej.EXAMPLE_RECORD]) # {qid: [p for each option, in option order]}
ej.load also takes a local model.ejpack (or a directory holding it). python examples/quickstart.py --weights 5ak3t/ej --revision v0.0.1 prints the distributions for ej.EXAMPLE_RECORD. Input and output format: the repository README.
One fit on one training pool (9,719 records, 7 source groups; sha256
ce1c1a6fecfd62a90317f6efc4f90fd5f9261becb7081fd00235a6f4d5ee9dbe). Details and counts: docs/data-card.md.
Amazon counterfactual is not used: its upstream licence is CC BY-NC 4.0 (one mirror declares CC BY 4.0).
Disclosure: Claude Haiku. The in-domain support tickets were written by Claude Haiku (labels by construction from the
generation spec, no human review). They are part of training; the corpus itself is not released, so the training pool
cannot be rebuilt from public data alone. Open licence questions on this and other inputs: NOTICE (Q1-Q4).
Not in the training data: Typed Decisions security_incidents (the zs_td suite), MASSIVE and the zs_wide workflows
(evaluation only).
- Suites (benchmarks/README.md; never trained on):- td: Typed Decisions test split, the 3 training workflows.
- zs_td: zs_td is ONE held-out workflow (security_incidents; dev 300 and final 100 records of the same workflow). The model code was kept on its dev accuracy after about 55 logged comparisons on it, so final zs_td estimates accuracy on further records of a workflow the model was selected on: it is neither unbiased nor unseen-workflow evidence. Its fit-to-fit SD (same code, other RNG) is ~.022 accuracy per fit (SD of a difference .031, 4 pairs), twice its record-cluster SE (.011); one workflow gives no between-workflow variance. Unseen-workflow evidence = zs_wide final macro_real.
- zs_massive: MASSIVE (en) intents with leak-free option sets. MASSIVE was never trained on, but its label space is not new: about 22% of its intent names (13 of 59 development option texts: 1 identical, 12 near) overlap CLINC150 / Banking77 option texts in the pool, and 4 final records are exact duplicates of pool records (removed at the next benchmark run).
- zs_wide: workflows never trained on: dev 154 workflows (111 from public sources: SNI tasks, SGD services, ABCD; 43 GLM-synthetic), final 147. The split is by task, not by dataset family: 53 of the 147 final workflows share a family with a dev workflow; pool-source tasks were dropped from the final set.
- tickets(in-domain, private) and- tickets_ood(held-out writing styles of the same generator: a style shift).
- Metrics: per-suite mean NLL of the gold option (headline: the mean over td, zs_td, zs_massive, tickets), micro
accuracy (zs_wide: group-macro over the real sources), ECE over 15 bins with its perfect-calibration floor, certified
automation (risk .10, δ .10; this suite only). Each cell has a 95% record-cluster CI. Scoring: ej.eval(benchmarks/scoring.pyfor rivals).
- Selection. Development numbers are selected: more than 55 logged comparisons were made on the development suites, so their numbers are optimistic. The final suites are read once per release; zs_td final is not unseen-workflow evidence (above).
- Leak-free option sets: for sampled option sets every option is equally likely to be the gold, so option frequency reveals nothing (a frequency rule reached .631 micro accuracy against chance .240 on a non-leak-free MASSIVE suite).
- Rivals and contamination flags: benchmarks/METHOD.md.
ej 0.0.1 (final suites, scored once, 2026-10-08: benchmarks/results/README.md). 95% record-cluster CIs; calibrated
= observed ECE15 at or below the 95th percentile of a perfectly calibrated model's ECE15 on that suite.
zs_wide final macro_real (group-macro accuracy over the real sources SNI, SGD and ABCD, 105 workflows; t interval over
workflows within sources) .419 [.379, .458]: the unseen-workflow number. GLM-synthetic workflows separately: .358
[.323, .394] (42 workflows). Mean NLL over td, zs_td, zs_massive and tickets: .7297. Seed variability of this code and
data (6 fits, development suites): between-seed SD .0105 in that metric, .019 in zs_wide macro_real, .037 in zs_td micro
accuracy; differences of that size between fits are not evidence of a change.
Against the rivals' sealed runs (paired CIs in benchmarks/results/README.md), ej 0.0.1 has lower micro accuracy than the
144M-0.8B rivals and the hosted Jev on zs_td and zs_massive (except Julia-1 on zs_massive) and than laya on td; no
calibration, latency or size ranking is claimed.
Micro accuracy / mean gold-label NLL on each final suite. Rival rows are their sealed runs on the same suites (2026-10-07);
the full table with ECE and paired 95% CIs of every difference is benchmarks/results/README.md.
¹ In-house ticket suites: ej trained on the tickets training split (home turf for ej, zero-shot for every rival).
² Jev's training data is undisclosed; its published Typed Decisions results use the same test split.
³ Likely contaminated: laya was trained on the Typed Decisions train split including security_incidents (so zs_td is
not zero-shot for laya); Julia-1 likely saw all TD workflows; OpenThai trains on MASSIVE intents.
Paired differences (ej minus rival, micro accuracy, 95% CI): vs Jev td −.015 [−.043, +.011], zs_td −.320, zs_massive −.125; vs laya td −.045 [−.065, −.023], zs_td −.344; vs gliclass-edge td +.355, zs_td −.062 [−.124, −.004], zs_massive +.287. On the 147-workflow unseen suite (zs_wide) no rival was run; ej's group-macro over real sources is .419. Latencies of the rivals were measured under a different protocol and are not compared here.
- Unseen workflows: low accuracy. Expect weak decisions on a new workflow until you measure it on labelled cases of
your own. A few labelled records help mainly by teaching the label prior (docs/adaptation.md).
- Fit-to-fit noise. Two fits of the same code that differ only in their random seed differ by about .03 zs_td micro accuracy (SD of a difference); single-fit differences of that size are not evidence (§4).
- Option-text tilt: on new workflows the option texts alone still tilt the predictions.
- More data did not simply help: adding broad text and dialogue data made td, tickets and zs_massive worse (zs_td unchanged within noise), and synthetic workflow data did not improve unseen-workflow accuracy; neither is in the training data.
- English only (base encoder and all data). Synthetic in-domain data: the tickets come from one model family (Claude Haiku) and have no human labels.
- Certified automation is per suite: a threshold certified on one workflow can fail on another (a td-certified threshold had error .771 on zs_wide). Certify on labelled rows of the workflow you automate.
- Python runtime only (torch + transformers, CPU). Latency is the Python reference on a 4-core CPU, not a phone.
- Float noise: batch composition, chunk size and thread count move probabilities by at most about 6e-7 (measured).
- One download: ej.load('5ak3t/ej', revision='v0.0.1')fetches onlymodel.ejpack, which holds the tokenizer and the encoder config; a local pack needs no network. A weights directory (the secondary format) instead fetches the base model from the Hugging Face Hub at revisionffb93f3bon first use.
- Reproducibility: a release is tied to its runtime (ej/_runtime, sha256 manifest checked at every load); a known pack is checked against its content digest inej.integrity.KNOWN_PACKS(and, when downloaded, its file sha256 inej.integrity.KNOWN_PACK_FILES).
- On-device triage and routing decisions (support tickets, intents, workflow checks) where a probability distribution is needed per question and uncertain cases are escalated to a slower System 2 (an on-device LLM or a cloud model).
- Ranking, gating and abstention using the returned probabilities, with thresholds certified on your own labelled data.
- Sole decision-maker for consequential decisions about people (credit, employment, health, legal, safety) without human review.
- Non-English text; generative tasks; free-text answers.
- New structured workflows without first measuring quality on labelled examples of that workflow (§6).
- Security-incident triage as a zero-shot claim (that workflow drove model selection).
- Transductive use across workflows, or on option sets that change between records: adapt/observeare for one workflow with a fixed option set (§1).
The weights are licensed under Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0). Adaptations must be shared under CC BY-SA 4.0 or a compatible licence. Suggested attribution:
ej 0.0.1 by Saket Bhushan, licensed under CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/). Derived from intfloat/e5-small-v2 (MIT; Wang et al., arXiv:2212.03533) and trained on Typed Decisions (Apache-2.0), Banking77 (CC BY 4.0), CLINC150 (CC BY 3.0), GoEmotions (Apache-2.0) and an in-house support-ticket corpus written with Claude Haiku (not released), with distillation from cross-encoder/nli-deberta-v3-xsmall (Apache-2.0). Full notices, creators and open questions: NOTICE.