Deterministic inference · Adapter training · Apache-2.0
typed-lm turns dense decoder models — Llama, Qwen2, Qwen3, Mistral, Gemma, Gemma2 and Gemma3 — into a typed semantic-routing API. Send a state and typed questions; receive booleans, choices and scores your code can branch on. No text generation, no parsing.
One forward pass means milliseconds, not seconds. On a single RTX 3070 with F16 weights, a full request — the shared prefill plus five batched question suffixes — is answered in tens to hundreds of milliseconds.
GPU (release, Qwen2.5-1.5B, F16, RTX 3070):
Adding a question adds a suffix to the same batched pass, not a new request, so latency grows with the prefix length — not with the number of questions.
CPU numbers (release, dense F32)
The recommended CPU mode is a GGUF Q4_K_M checkpoint with the mkl feature.
The session prefix cache skips the prefill entirely for repeated states. More in benchmarks.
A large language model answers by generating text token by token. When your software needs a judgment it can branch on, that creates a mismatch: you prompt, you parse, you validate — and you still get a string. typed-lm removes the mismatch. It runs the model once, reads the logits at a single decision position, and returns a typed value with a calibrated distribution.
flowchart LR
client["Client"]:::neutral
request["state + questions"]:::primary
subgraph model["typed-lm-serve"]
direction TB
prefill["shared prefill"]:::accent
batch["batched decision positions"]:::accent
end
answers["typed answers<br/>noul · choice · score"]:::success
code["your code<br/>branch · sort · route"]:::success
client --> request --> prefill --> batch --> answers --> code
classDef primary fill:#ede9fe,stroke:#7c3aed,color:#3b0764,stroke-width:1.5px
classDef accent fill:#dbeafe,stroke:#2563eb,color:#0c4a6e,stroke-width:1.5px
classDef success fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:1.5px
classDef neutral fill:#f4f4f5,stroke:#a1a1aa,color:#18181b,stroke-width:1.5px
All three can be combined in a single request, and each question is evaluated independently against the same state.
flowchart LR
state["state"]:::neutral
noul["noul question"]:::primary
choice["choice question"]:::accent
score["score question"]:::success
answers["answers map"]:::success
state --> noul --> answers
state --> choice --> answers
state --> score --> answers
classDef primary fill:#ede9fe,stroke:#7c3aed,color:#3b0764,stroke-width:1.5px
classDef accent fill:#dbeafe,stroke:#2563eb,color:#0c4a6e,stroke-width:1.5px
classDef success fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:1.5px
classDef neutral fill:#f4f4f5,stroke:#a1a1aa,color:#18181b,stroke-width:1.5px
typed-lm is not just an inference server — it ships a trainer that turns a general-purpose checkpoint into a specialist for your decisions. It optimizes the cross-entropy at the decision position, the exact position the server reads, so what you train is what you serve.
Why train with typed-lm?
- One objective, end to end — the training loss is the serving decision, so there is no train/serve skew.
- Cheap specialization — LoRA/QLoRA store only the adapter tensors; the base is never duplicated.
- Your labels, your thresholds — confidence is calibrated on your data.
- Quantize what you train — FP8/FP4 PTQ and full/from-scratch checkpoints are served by the same binary, with no merge step for complete checkpoints.
flowchart LR
dataset["dataset<br/>state + questions + answer"]:::neutral
checkpoint["base checkpoint"]:::accent
config["run configuration<br/>CLI or TOML"]:::warning
train["train<br/>lora · qlora · full · from-scratch"]:::primary
artifact["artifact<br/>adapter or checkpoint"]:::success
quantize["quantize<br/>fp8 · fp4"]:::accent
serve["typed-lm-serve"]:::success
dataset --> train
checkpoint --> train
config --> train
train --> artifact --> serve
artifact --> quantize --> serve
classDef primary fill:#ede9fe,stroke:#7c3aed,color:#3b0764,stroke-width:1.5px
classDef accent fill:#dbeafe,stroke:#2563eb,color:#0c4a6e,stroke-width:1.5px
classDef success fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:1.5px
classDef warning fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:1.5px
classDef neutral fill:#f4f4f5,stroke:#a1a1aa,color:#18181b,stroke-width:1.5px
# Train a LoRA adapter over a frozen checkpoint.
typed-lm-trainer train \
--model-id /path/to/local/checkpoint \
--dataset resources/dataset.jsonl \
--output-directory output/train \
--method lora --epochs 3 --batch-size 4 --learning-rate 1e-4
# Merge the adapter and quantize to FP8.
typed-lm-trainer quantize \
--model-id /path/to/local/checkpoint \
--adapter-directory output/train \
--quantization fp8 --output-directory output/quantizedFull tutorial: training.
The fastest path — no toolchain, just an image. The server image pulls the model
on first startup and listens on 8080:
# Server. Pass an HF_TOKEN for gated models and mount a context file if you have one.
docker run --rm -p 8080:8080 \
-e HF_TOKEN=<hugging-face-token> \
-v "$PWD/resources/memory.md:/etc/typed-lm/memory.md:ro" \
-e CONTEXT_PATH=/etc/typed-lm/memory.md \
ghcr.io/neurono-ml/typed-lm-serve:0.1.1
# Ask three typed questions in one call.
curl -s http://127.0.0.1:8080/v1/systemone \
-H 'Content-Type: application/json' \
-d @examples/request_mixed.jsonThe trainer runs the same way, with the artifacts directory mounted so the outputs survive the container:
# Train a LoRA adapter; /work holds the checkpoint, dataset and outputs.
docker run --rm -v "$PWD:/work" -w /work \
-e HF_TOKEN=<hugging-face-token> \
ghcr.io/neurono-ml/typed-lm-trainer:0.1.1 train \
--model-id /work/checkpoint \
--dataset /work/resources/dataset.jsonl \
--output-directory /work/output/train \
--method lora --epochs 3 --batch-size 4 --learning-rate 1e-4For a GPU, use the :cuda image (it includes the CUDA runtime libraries) and
pass --gpus all; the host only needs the NVIDIA driver and the container
toolkit:
# Server on GPU.
docker run --rm --gpus all -p 8080:8080 \
-e HF_TOKEN=<hugging-face-token> \
ghcr.io/neurono-ml/typed-lm-serve:cuda
# Trainer on GPU.
docker run --rm --gpus all -v "$PWD:/work" -w /work \
-e HF_TOKEN=<hugging-face-token> \
ghcr.io/neurono-ml/typed-lm-trainer:cuda train \
--model-id /work/checkpoint \
--dataset /work/resources/dataset.jsonl \
--output-directory /work/output/train \
--method lora --device cuda --epochs 3 --batch-size 4 --learning-rate 1e-4# Install (CPU build; add --features cuda or --features metal for a GPU).
cargo install typed-lm-serve typed-lm-trainer
# Start the server (downloads the default model on first startup).
typed-lm-serve --context-path resources/memory.md
# Ask three typed questions in one call.
curl -s http://127.0.0.1:8080/v1/systemone \
-H 'Content-Type: application/json' \
-d @examples/request_mixed.jsonRequest
{
"model": "typed-lm",
"state": "Order #7710 arrived with a smashed box and a cracked vase inside. Delivery was 3 days ago and the customer asks what to do next.",
"questions": {
"refund_eligible": {
"type": "noul",
"instructions": "The customer is eligible for a full refund under the store policy."
},
"responsible_department": {
"type": "choice",
"instructions": "Which department should handle this case?",
"criteria": {
"billing": "Double charges and payment errors",
"logistics": "Damaged, lost, or late shipments",
"product_support": "Defective-item troubleshooting, replacements, and setup help"
}
},
"urgency": {
"type": "score",
"instructions": "How urgent is this case?",
"criteria": ["Routine", "Urgent", "Emergency"]
}
}
}Response
{
"model": "typed-lm",
"answers": {
"refund_eligible": { "type": "noul", "noul": 0.87 },
"responsible_department": {
"type": "choice",
"choice": "logistics",
"probabilities": { "billing": 0.05, "logistics": 0.9, "product_support": 0.05 },
"confidence": 0.85
},
"urgency": {
"type": "score",
"score": 1.2,
"legend": { "0": "Routine", "1": "Urgent", "2": "Emergency" },
"probabilities": { "0": 0.2, "1": 0.4, "2": 0.4 },
"confidence": 0.2
}
},
"usage": { "input_tokens": 512, "output_tokens": 4 }
}Full walkthrough: quickstart.
Detected automatically from model_type in config.json.
flowchart TB
config["config.json model_type"]:::neutral
dense{"dense family?"}:::warning
family["llama · qwen2 · qwen3<br/>mistral · gemma · gemma2 · gemma3"]:::success
moe["mixtral · qwen3_moe<br/>deepseek_v2 · deepseek_v3"]:::danger
served["served"]:::success
rejected["rejected"]:::danger
config --> dense
dense -- "yes" --> family --> served
dense -- "no" --> moe --> rejected
classDef success fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:1.5px
classDef danger fill:#fee2e2,stroke:#dc2626,color:#7f1d1d,stroke-width:1.5px
classDef warning fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:1.5px
classDef neutral fill:#f4f4f5,stroke:#a1a1aa,color:#18181b,stroke-width:1.5px
Dense safetensors, PyTorch (.pth/.bin) and NumPy (.npz) checkpoints of any
of the seven families are served. GGUF-quantized serving is Qwen2-only.
Mixture-of-Experts and multi-head-latent-attention families are rejected at load
time. FP8 and FP4 artifacts are dequantized on load; GPTQ/AWQ are rejected.
cargo build --workspace
cargo run -p typed-lm-serve -- --help
cargo run -p typed-lm-trainer -- --helpPrebuilt images for both binaries are published to the GitHub Container Registry
on every release. CPU images carry latest and the version; CUDA images add a
-cuda suffix (and the cuda tag):
flowchart LR
host["host"]:::neutral
gpu{"NVIDIA GPU<br/>+ container toolkit?"}:::warning
cpu["typed-lm-serve:0.1.1<br/>typed-lm-trainer:0.1.1<br/>(latest too)"]:::accent
cuda["typed-lm-serve:cuda<br/>typed-lm-trainer:cuda"]:::success
runcpu["docker run -p 8080:8080"]:::accent
runcuda["docker run --gpus all<br/>--model-dtype auto"]:::success
serve["typed answers"]:::primary
host --> gpu
gpu -- "no" --> cpu --> runcpu --> serve
gpu -- "yes" --> cuda --> runcuda --> serve
classDef primary fill:#ede9fe,stroke:#7c3aed,color:#3b0764,stroke-width:1.5px
classDef accent fill:#dbeafe,stroke:#2563eb,color:#0c4a6e,stroke-width:1.5px
classDef success fill:#d1fae5,stroke:#059669,color:#064e3b,stroke-width:1.5px
classDef warning fill:#fef3c7,stroke:#d97706,color:#78350f,stroke-width:1.5px
classDef neutral fill:#f4f4f5,stroke:#a1a1aa,color:#18181b,stroke-width:1.5px
# CPU
docker pull ghcr.io/neurono-ml/typed-lm-serve:latest
docker pull ghcr.io/neurono-ml/typed-lm-serve:0.1.1
# CUDA (GPU)
docker pull ghcr.io/neurono-ml/typed-lm-serve:cuda
docker pull ghcr.io/neurono-ml/typed-lm-serve:0.1.1-cudaThe server listens on 8080; pass an HF_TOKEN for gated models, and mount a
context file and the model cache:
docker run --rm -p 8080:8080 \
-e HF_TOKEN=<hugging-face-token> \
-v typed-lm-cache:/root/.cache/huggingface \
-v "$PWD/resources/memory.md:/etc/typed-lm/memory.md:ro" \
-e CONTEXT_PATH=/etc/typed-lm/memory.md \
ghcr.io/neurono-ml/typed-lm-serve:0.1.1The CUDA images bundle the runtime libraries candle loads (cudart, cublas,
curand, nvrtc); the host only needs the NVIDIA driver and the container
toolkit. Select F16 weights automatically with --model-dtype auto:
docker run --rm --gpus all -p 8080:8080 \
-e HF_TOKEN=<hugging-face-token> \
-e MODEL_DTYPE=auto \
-v typed-lm-cache:/root/.cache/huggingface \
ghcr.io/neurono-ml/typed-lm-serve:cudaTo build the CUDA image from source instead (the release pipeline does this automatically), use the multi-stage Dockerfile; the compute capability can be tuned for the target GPU:
docker build -f docker/Dockerfile.serve-cuda \
--build-arg CUDA_COMPUTE_CAP=80 -t typed-lm-serve:cuda .Each release attaches binaries for Linux x86_64 (CPU/CUDA) and macOS arm64 (Metal):
Full CPU and GPU latency tables, the acceleration features and the session-cache gain are in benchmarks.
POST /v1/systemone, GET /v1/models, GET /health, GET /health/live.
Invalid bodies return 422, unknown models 404, inference failures 500, all
with the {"error": {"message": "..."}} envelope.
cargo test --workspace # unit + integration, no download
cargo test --workspace -- --ignored --nocapture # live tests (real weights)
cargo clippy --workspace --all-targets
cargo fmt --checkcargo test --workspace includes a weight-free binary E2E
(train → quantize → serve over HTTP). Live tests marked #[ignore] need
real weights and a GPU for the training cases; they never run in CI.
The complete guide is published at https://neurono-ml.github.io/typed-lm/:
For AI assistants, the site exposes an index at https://neurono-ml.github.io/typed-lm/llms.txt.
Contributions are welcome — code, docs, datasets and prompts alike. The project
rules live in AGENTS.md.
cargo fmt --all --check
cargo clippy --workspace --all-targets -- -D warnings
cargo test --workspaceApache-2.0. See LICENSE.