Give it a state and a set of questions — choice, score, or yes-no. It returns one probability per option, in a single forward pass.
No tokens and no free text, so it cannot answer outside your candidate set. It runs as a single local executable, which keeps inference inside your perimeter.
Accuracy against what a decision costs. A workflow is five typed decisions about one record. Competitor accuracies and prices are as published; ours is measured over all 2,000 test decisions and priced from a laptop’s measured throughput plus its amortised system cost. The full basis is in the report.
AT0M will ship as a single self-contained Rust binary — no Python at inference, no service in the path, nothing to rent. That is the deployment, and it will be the only thing that ends up in your perimeter.
The hosted endpoint exists for one reason: so you can check every number on this page before you download anything.
v0 is the evaluation endpoint; v1 is the executable release — the hosted service is a way to try the model, not a way to run it. Every figure on this page was measured against v0, the evaluation build, so they are the floor rather than the ceiling: v1 is expected to be more accurate, and we will republish every number on this page when it ships.
One configuration across every public test set — no per-dataset tuning or prompt tweaks. 38,862 decisions, scored in one sweep.
Laya figures, and Jev's on the classification sets, are as reported in the Laya write-up. Jev's typed-decisions figure is DecisionEval's independent rerun. On calibration, TypeSafe publishes no figure at all, so Jev's range is assembled from four independent studies, none temperature-fitted.
The accuracy number is near the top but sits inside admitted label noise, so we don't hang the pitch on it. The claims we'd stand behind under scrutiny are architectural:
And the parts we wouldn't: it compares a value to a threshold well, but not one field of a record against another, so compute those predicates in code. Calibration decays off-distribution — 0.037 on spam it has seen, 0.171 on phishing it hasn't. Both are measured below.
Every figure on this page came from a public test split scored through this endpoint. Point the same benchmark at it and check. One POST carries a record and as many questions as you like, all answered in the same pass — no SDK, no client library.
We'll help you scope a local-training or private-deployment evaluation against your own decisions — and share the full technical report.
Request an evaluation Get the technical reportA comparative evidence inventory across AT0M, Jev 1.13.0 and Laya, carried from the technical paper. It is not a harmonized three-model benchmark: figures were obtained under different conditions, and each is tagged with its source and how it was measured.
Sources: [S1] Pienomial AT0M benchmark report v1.0, 27 Sep 2026 · [L1/L2] Convai Laya model cards · [I1] DecisionEval Jev 1.13.0 independent rerun, 20 Sep 2026 · [J1/J2] TypeSafe documentation · [B1] public typed-decisions leaderboard. Latencies are not hardware- or network-normalized; pricing and availability may change after the report date. "Not disclosed" means the cited sources don't establish the item — not that it's unavailable.
V1 is the single binary you run on your own hardware. Leave an address and we'll tell you when it ships — nothing else, and no one else gets it.