Fast

Decisions in milliseconds.

A decision model answers in a single forward pass, with no token-by-token generation. On your own GPU, a five-question request to Laya takes about 10 ms, end to end through the HTTP API.

8–10ms

Laya on Ollaya

RTX 4090, five questions, end to end

236–276ms

TypeSafe Jev

Hosted API, median request

Drop-in compatible

Speaks TypeSafe's API.

Ollaya serves /v1/systemone and /v1/models with TypeSafe's request and response shapes. The official TypeSafe Python SDK 0.7.1 works unchanged against a local server.

Request

# Point the TypeSafe SDK at Ollaya

export TYPESAFE_BASE_URL=http://localhost:11435

export TYPESAFE_API_KEY=local # any value works

export TYPESAFE_DEFAULT_MODEL=laya

# …or call the compatible endpoint directly

curl http://localhost:11435/v1/systemone -d '{

"model": "laya",

"state": "Can I get an invoice for last month?",

"questions": {

"intent": {

"type": "choice",

"instructions": "What does the customer want?",

"criteria": {

"invoice": "Needs an invoice or receipt",

"refund": "Wants money back",

"other": "Anything else"

}

}

}

}'Response

{

"model": "laya:en",

"answers": {

"intent": {

"type": "choice",

"choice": "invoice",

"confidence": 0.9547,

"probabilities": {

"invoice": 0.9698,

"refund": 0.0172,

"other": 0.013

}

}

},

"usage": {

"input_tokens": 43,

"output_tokens": 0

}

}Open models

Open weights, ready to pull.

Start with Laya from Convai Innovations: an English model, a 100+ language model, a model fine-tuned for typed decisions, and a router that picks for you.

- layaOpen decision models from Convai Innovations. Typed, calibrated answers to choice, score and yes/no questions in a single forward pass, in English and 100+ languages.322m · 421m

- deciderDecoder decision models by Mapika on Qwen3.5: the answer is read from option-letter logits in one forward pass. The most accurate open decision model Ollaya ships.0.75b · 1.9b

- nliZero-shot classifiers by Moritz Laurer: every option becomes a hypothesis scored for entailment. The most accurate encoder model on typed decisions in our tests.396m · 435m

- gliclassInstruction-following zero-shot classifier by Knowledgator: all options of a question are scored in one pass, so cost barely grows with the number of options.439m

- qwen3guardSafety guard by the Qwen team: is a text safe, controversial or unsafe, and which unsafe category? It answers its own built-in questions, in 119 languages, in one forward pass.0.6b

- kevDecision models by Jared Palmer: a LoRA on a Qwen3.5 base plus a pointer head that scores every option at its own span, in one forward pass per question. Calibrated with Kev's own temperature.0.76b

- vonDecision model by Victor Hugo Panisa on ModernBERT-large: every option is scored at its own marker, all options of a question in one pass, with an input-conditioned calibration. 8k-token context.395m

Your data stays yours

Private by default.

Tickets, emails and user messages are often the most sensitive data you have. With Ollaya they are scored where they already live.

- Local- Runs on your machine with ONNX Runtime, on the CPU or an NVIDIA GPU. The server listens on 127.0.0.1 by default.

- Open weights- Weights come from their authors’ Hugging Face repositories, pinned to a commit and checked against sha256. Ollaya never re-hosts them, and the runtime is Apache-2.0.

- No per-token fees- Run as many decisions as your hardware can handle. No metering and no API bill.

- Calibrated- Probabilities you can put thresholds on. Each model ships its own calibration, and a Modelfile refits it on your labelled data.

Platforms

Runs where you work.

A desktop app and a command line for macOS, Windows and Linux, and a Docker image for servers. Every model runs on the CPU; an NVIDIA GPU on Linux, in WSL 2 or in Docker takes a request down to milliseconds.

Get up and running in minutes.

One binary, one command: ollaya run laya.

macOS, Windows, Linux and Docker · Apache-2.0 · GitHub