One forward pass reads a document and returns a calibrated probability for every option of every question you ask. Nothing is generated.

Rene-1 is a decision model. You send a document and a set of typed questions: pick one option, answer yes or no, or place the document on a scale. It reads the document and every question in one pass, each question on its own, and answers each with a probability for every option you declared. There is no generated text to parse. The weights are stored in FP8, so the whole model runs on one GPU.

Install PyTorch first, built for your CUDA version (torch>=2.14.0, from pytorch.org), then:

pip install "transformers>=5.17.0" "compressed-tensors>=0.19.0" "accelerate>=1.15.0" "safetensors>=0.8.0" "pillow>=12.3.0"

Rene-1 needs an NVIDIA GPU that runs FP8 and holds its weights, such as those under Hardware, below. On any other device it refuses to load, and says why.

Every classifier gets asked the hot dog question sooner or later. One request asks three typed questions about a dish: yes or no, which cuisine, and how messy it is to eat.

from transformers import AutoModel

model = AutoModel.from_pretrained("salfatigroup/rene-1-31b-fp8", trust_remote_code=True, device_map="auto")

state = {"dish": "A grilled sausage in a split bun with mustard"}

questions = {

"is_hotdog": {

"type": "yesno",

"instructions": "Is the food in `dish` a hot dog?",

"criteria": {

"true": "A sausage served in a sliced bun.",

"false": "Anything else."

}

},

"cuisine": {

"type": "choice",

"instructions": "Which cuisine does `dish` belong to?",

"criteria": {

"american": "US street food or diner food.",

"german": "German cooking.",

"other": "Anything else."

}

},

"messiness": {

"type": "score",

"instructions": "How messy is `dish` to eat by hand?",

"criteria": [

"clean",

"a little messy",

"very messy"

]

}

}

answers = model.decide(state, questions)

print(answers["is_hotdog"])

The response to the hotdog request

The response:

{

"model": "rene-1-31b-fp8",

"answers": {

"is_hotdog": {

"type": "yesno",

"probability": 0.94

},

"cuisine": {

"type": "choice",

"choice": "american",

"confidence": 0.63,

"probabilities": {

"american": 0.75,

"german": 0.22,

"other": 0.03

}

},

"messiness": {

"type": "score",

"score": 1.14,

"confidence": 0.63,

"legend": {

"0": "clean",

"1": "a little messy",

"2": "very messy"

},

"probabilities": {

"0": 0.06,

"1": 0.75,

"2": 0.19

}

}

},

"usage": {

"input_tokens": 354,

"output_tokens": 0

}

}

Each response carries:

- typeon every answer.

- probabilityfor a yes or no question: the probability of- true.

- probabilitiesfor a choice, keyed by option, and for a score, keyed by level number as a string, "0" for the lowest. They sit on a two-decimal grid and sum to 1.

- legendfor a score: each level number with its text.

- choice, the most likely option. Its- confidenceis (K x top probability - 1) / (K - 1) over K options: 0 for an even split, 1 for certainty.

- score, the expected level, counting levels from 0. Its- confidenceis 1 minus the expected distance from the most likely level, scaled so an even spread scores 0.

- usage.output_tokens, which is 0 because nothing is generated.

trust_remote_code=True runs the model code in this repository (modeling_rene.py and the files next to it). Read it first, and pin the commit you read with revision= in production.

- model.decide(state, questions)returns the answers, keyed by your question names.

- model.decide_batch(requests)answers up to 64 requests, 65,536 packed tokens in total, in one call. On a GPU a batched answer can differ from the single call's by at most 0.01 on the two-decimal grid.

- model.respond(request)and- model.respond_batch(requests)return the whole response object,- modeland- usageincluded.

- model.render(state, questions)shows the exact sequence the model reads.

- The tokenizer: AutoTokenizer.from_pretrained("salfatigroup/rene-1-31b-fp8", subfolder="trunk"), ormodel.tokenizer.

- At load it warns once if an installed package differs from the versions it was tested with, listed in release.json.

A field of state can hold an image: a PIL image or the file's bytes in Python, or the JSON form below for model.respond (PNG, JPEG or WebP). Refer to it by name in backticks, like any other field. Up to 4 images per request; each adds about 280 input tokens. URLs and file paths are refused, so the code never fetches anything. Image input is experimental: Rene-1 was trained on text only, image answers have not been evaluated, and the model says so once when it serves one.

from PIL import Image

questions = {

"is_hotdog": {

"type": "yesno",

"instructions": "Is the food in `photo` a hot dog?"

},

"dish": {

"type": "choice",

"instructions": "Which dish is in `photo`?",

"criteria": {

"hot_dog": "A sausage in a sliced bun.",

"corn_dog": "A battered sausage on a stick.",

"sandwich": "Filling between slices of bread.",

"other": "Anything else."

}

}

}

answers = model.decide({"photo": Image.open("lunch.png")}, questions)

print(answers["is_hotdog"])

The same request as JSON:

{

"state": {

"photo": {

"type": "image",

"base64": "<the file, base64>",

"media_type": "image/png"

}

},

"questions": {

"is_hotdog": {

"type": "yesno",

"instructions": "Is the food in `photo` a hot dog?"

},

"dish": {

"type": "choice",

"instructions": "Which dish is in `photo`?",

"criteria": {

"hot_dog": "A sausage in a sliced bun.",

"corn_dog": "A battered sausage on a stick.",

"sandwich": "Filling between slices of bread.",

"other": "Anything else."

}

}

}

}

The response for `hot-dog-drawing.png`, one of the drawings in the video above

{

"model": "rene-1-31b-fp8",

"answers": {

"is_hotdog": {

"type": "yesno",

"probability": 0.78

},

"dish": {

"type": "choice",

"choice": "hot_dog",

"confidence": 0.94,

"probabilities": {

"hot_dog": 0.96,

"corn_dog": 0.0,

"sandwich": 0.01,

"other": 0.03

}

}

},

"usage": {

"input_tokens": 590,

"output_tokens": 0

}

}

The response for `corn-dog-drawing.png`, one of the drawings in the video above

{

"model": "rene-1-31b-fp8",

"answers": {

"is_hotdog": {

"type": "yesno",

"probability": 0.34

},

"dish": {

"type": "choice",

"choice": "other",

"confidence": 0.31,

"probabilities": {

"hot_dog": 0.1,

"corn_dog": 0.4,

"sandwich": 0.01,

"other": 0.49

}

}

},

"usage": {

"input_tokens": 580,

"output_tokens": 0

}

}

The same model as a transformers pipeline, which returns the whole response object:

from transformers import pipeline

rene = pipeline(model="salfatigroup/rene-1-31b-fp8", trust_remote_code=True, device_map="auto")

response = rene({"state": state, "questions": questions})

Pick the queue, flag the tickets that cannot wait, and read the customer's mood, in one pass.

The request and the response

The request:

{

"state": {

"subject": "Charged twice for the annual plan",

"body": "I moved to the annual plan yesterday and my card shows the same charge twice. I only want one plan. Please refund the duplicate before it posts.",

"plan": "business"

},

"questions": {

"queue": {

"type": "choice",

"instructions": "Which team should handle the ticket in `subject` and `body`?",

"criteria": {

"billing": "Charges, refunds, invoices and plan changes.",

"technical": "Errors, outages, bugs and integrations.",

"account": "Sign-in, access, security and profile settings.",

"sales": "Questions from people who have not bought yet."

}

},

"urgent": {

"type": "yesno",

"instructions": "Does the ticket in `body` need a reply within the hour?",

"criteria": {

"true": "Money taken in error, a security problem, or a service that is down.",

"false": "Anything that can wait for the normal queue."

}

},

"frustration": {

"type": "score",

"instructions": "How frustrated is the customer who wrote `body`?",

"criteria": [

"calm",

"annoyed",

"angry"

]

}

}

}

The response:

{

"model": "rene-1-31b-fp8",

"answers": {

"queue": {

"type": "choice",

"choice": "billing",

"confidence": 1.0,

"probabilities": {

"billing": 1.0,

"technical": 0.0,

"account": 0.0,

"sales": 0.0

}

},

"urgent": {

"type": "yesno",

"probability": 0.62

},

"frustration": {

"type": "score",

"score": 0.54,

"confidence": 0.19,

"legend": {

"0": "calm",

"1": "annoyed",

"2": "angry"

},

"probabilities": {

"0": 0.56,

"1": 0.34,

"2": 0.1

}

}

},

"usage": {

"input_tokens": 435,

"output_tokens": 0

}

}

Classify the clause, check whether its right runs both ways, and place its risk on a four-level scale.

The request and the response

The request:

{

"state": {

"clause": "Either party may end this contract for convenience with thirty days' written notice. On termination the customer pays all fees accrued up to the termination date, and the provider refunds any prepaid fees for the period after it."

},

"questions": {

"clause_type": {

"type": "choice",

"instructions": "What kind of clause is `clause`?",

"criteria": {

"termination": "When and how the contract can end.",

"liability": "Caps on, or exclusions of, damages.",

"payment": "Fees, invoicing and payment terms.",

"confidentiality": "Duties to keep information secret.",

"other": "Anything else."

}

},

"mutual": {

"type": "yesno",

"instructions": "Can both parties use the right that `clause` grants?"

},

"customer_risk": {

"type": "score",

"instructions": "How much risk does `clause` put on the customer?",

"criteria": [

"low",

"moderate",

"high",

"severe"

]

}

}

}

The response:

{

"model": "rene-1-31b-fp8",

"answers": {

"clause_type": {

"type": "choice",

"choice": "termination",

"confidence": 0.97,

"probabilities": {

"termination": 0.98,

"liability": 0.0,

"payment": 0.01,

"confidentiality": 0.0,

"other": 0.01

}

},

"mutual": {

"type": "yesno",

"probability": 0.85

},

"customer_risk": {

"type": "score",

"score": 0.38,

"confidence": 0.62,

"legend": {

"0": "low",

"1": "moderate",

"2": "high",

"3": "severe"

},

"probabilities": {

"0": 0.77,

"1": 0.13,

"2": 0.05,

"3": 0.05

}

}

},

"usage": {

"input_tokens": 401,

"output_tokens": 0

}

}

The policy is part of the request, so changing a rule needs no retraining.

The request and the response

The request:

{

"state": {

"policy": "Harassment: posts must not insult, threaten or demean a person or a group. Criticism of ideas, products and the public actions of public figures is allowed.",

"post": "This update broke my workflow again. Whoever shipped this release has clearly never used the product."

},

"questions": {

"breaks_policy": {

"type": "yesno",

"instructions": "Does `post` break `policy`?"

},

"content": {

"type": "choice",

"instructions": "What is `post` mainly doing?",

"criteria": {

"product_criticism": "Criticising a product, a release or a decision.",

"personal_attack": "Insulting or demeaning a person or a group.",

"threat": "Threatening harm.",

"spam": "Advertising or repeated content.",

"other": "Anything else."

}

},

"severity": {

"type": "score",

"instructions": "If `post` breaks `policy`, how serious is it?",

"criteria": [

"not a violation",

"mild",

"serious"

]

}

}

}

The response:

{

"model": "rene-1-31b-fp8",

"answers": {

"breaks_policy": {

"type": "yesno",

"probability": 0.25

},

"content": {

"type": "choice",

"choice": "product_criticism",

"confidence": 0.55,

"probabilities": {

"product_criticism": 0.64,

"personal_attack": 0.31,

"threat": 0.01,

"spam": 0.01,

"other": 0.03

}

},

"severity": {

"type": "score",

"score": 0.46,

"confidence": 0.31,

"legend": {

"0": "not a violation",

"1": "mild",

"2": "serious"

},

"probabilities": {

"0": 0.61,

"1": 0.33,

"2": 0.06

}

}

},

"usage": {

"input_tokens": 412,

"output_tokens": 0

}

}

- stateis a string, an object or a list. Instructions refer to its fields by name in backticks.

- Questions are keyed by your own names, and answers come back under the same keys.

- A choice question with a single option is answered with certainty; the model does not read it.

- An unknown field in a request or a question is refused, not ignored, so a typo fails loudly.

- Each question sees the document and itself, never the other questions. The document plus any one question can hold 32,512 tokens, and a whole request 65,280, counted on the text you send (objects as compact JSON).

- A request the model refuses raises an error that says why.

The Decision Index is a public panel of 40 decision benchmarks in 5 areas. Its headline, balanced skill, corrects each benchmark for chance and weighs the areas equally. We rebuilt 37 of the 40 benchmarks from their public sources, 91.81 of its 100 points: 28 with the index's own kit (apolinario/decision-index at 8e12d6d7); BANKING77 and Habermas with the kit's own code, run on source files its pinned copy does not hold; and the 7 benchmarks new in Decision Index 0.2 (MMLU-Pro, BBH, RAGTruth, PhishNChips, HoVer, When2Call and New Yorker) with our own code. We evaluated Rene-1 on those rows at temperature 1, with no calibration set, on NVIDIA B200. The other rows are the index's own board, the maintainers' runs, averaged over the same 37 benchmarks; Rene-1 has not been scored by them. The edition is pinned to Decision Index 0.2, the board of 2026-09-24. The index has since published Decision Index 0.2.1 (board of 2026-09-27), a different panel: 38 benchmarks, with SGD and RouterBench out of the index; new area weights; 13 benchmarks weighted 1.2 inside their area; and changed scoring, chance or row counts on MuSR, HLE, ACOS, RAGTruth, BRIGHT, ToolRet and Home appliances. This page uses Decision Index 0.2 only and never mixes the two.

Not evaluated (3):

- ToolRet: its 0.2 row selection is not published, so an evaluation would not follow the 0.2 protocol.

- BRIGHT: its 0.2 row selection is not published, so an evaluation would not follow the 0.2 protocol.

- HLE: its rows need an access request on the Hub, which we did not make.

- On the same 37 benchmarks, Rene-1 places 1 of 51 against all 50 rows of the Decision Index 0.2 board.

Every column averages skills over the benchmarks evaluated here, so the board rows are on the same 37 as the first.

At temperature 1 and with no calibration set, the median calibration error (ECE, 10 equal-mass bins) across the 37 benchmarks is 0.046. The last column below gives it per benchmark.

Skill is each benchmark's own score corrected for chance, (score - chance) / (1 - chance), from 0 at chance to 100 for a perfect answer. A score below chance is floored at 0 and marked (floored). The best score on each row is in bold. ForecastBench, where a lower score is better, enters as clip((0.25 - Brier) / (0.25 - 0)) x coverage, so a higher skill is better there too.

JevBench at commit 1bcc55e, run on its 231 public items (48 easy, 72 standard, 111 hard) and scored with its own v1.3.0 formulas, the last it defines on the public items alone. Its judge tier is not public, so the benchmark's own rule spreads that weight over the other tiers. The benchmark's own board also counts sealed items that only its maintainers run, so its board rows are not set beside this local run.

- On the 111 hard items Rene-1 is right on 77.5%. Its weakest kinds of hard item are date and number arithmetic (2 of 15 right) and long policy documents (12 of 19 right).

- Speed follows the benchmark's rule for a self-hosted GPU: it times the standard and judge items one request at a time. The judge items are not public, so Rene-1 is timed here on the 72 standard items, on NVIDIA B200, and each latency is multiplied by 2 with 0.15 s added before scoring. Measured: p50 97.6 ms and p95 102.9 ms.

- Cost needs a public per-token price, which this page does not set, so Cost and the combined Score are left out.

Measured on 1x NVIDIA B200 (NVIDIA B200, driver 595.91.07) with torch 2.14.0+cu130 (CUDA 13.0), transformers 5.17.0 and compressed-tensors 0.19.0. Each row times model.decide end to end, as a caller sees it: request checks, tokenization, the one forward pass and the answer arithmetic, synchronised with the GPU, after warm-up calls. p95 is the nearest rank. The weights take 33.3 GB on disk; peak GPU memory in these runs was 39.3 GB allocated and 44.5 GB reserved. Loading took 12.9 s.

Throughput runs model.decide_batch on distinct requests and divides by the median batch time.

FP8 matrix multiplies need compute capability 8.9 or newer: Ada Lovelace, Hopper or Blackwell. Rene-1 checks the GPU when it loads and refuses anything else, rather than run a slower copy that was never measured. At 44.5 GB peak, one GPU holds the model: L40S (48 GB), H100 (80 GB), H200 (141 GB) or B200 (180 GB). L4 (24 GB) has the compute capability but not the memory. A request at the token limit (62,976 input tokens) peaked at 67.9 GB, so requests that long need H100 (80 GB), H200 (141 GB) or B200 (180 GB).

- Classifying, routing and triaging text against labels you define in each request.

- Many questions about one document in a single pass: checks on a contract, a ticket, a transcript or a record.

- Grading on an ordinal scale, with the uncertainty attached.

- Thresholds and review queues: act on the confident answers, send the uncertain ones to a person.

- The sole basis for a decision about a person's access to credit, work, housing, healthcare, education, legal status or benefits. Keep a person in the loop, and test on your own data first.

- Free text of any kind: summaries, explanations, chat. Rene-1 writes nothing.

- Multi-step arithmetic, date arithmetic and long chains of reasoning.

- Security and fraud triage without your own evaluation.

- Languages other than English. It was not evaluated on them.

- Images. The model code accepts an image inside stateas an experimental input, but the decision layer was trained on text only and image answers have not been evaluated; the model warns once when it serves one.

- Its weakest benchmarks by skill are ACOS (13.7), ChessBench (14.8) and SGD (18.3).

- The board's leading row scores higher than Rene-1 on 8 of the 37 benchmarks evaluated; the table under Results names them.

- Its calibration is weakest on PhishNChips (calibration error 0.527), SGD (calibration error 0.423) and POP909 (calibration error 0.372). Check its confidence on data like these before you threshold it.

- On the hard items of the third-party benchmark in Results it is right on 77.5%; it misses most on date and number arithmetic and long policy documents.

- It gives no reasons. A probability is not an explanation.

- It has not been tested for sensitivity to the order or the names of the options.

- Speed depends on the GPU and on the shape of the request. Measure yours.

Rene-1 is a 31B transformer fine-tuned in full from open weights: the language model was trained for one epoch on 396,500 rows from 56 task families under a log-score loss, to put its probability on the right option. The trained weights are stored in compressed-tensors FP8_DYNAMIC format: E4M3 weights with one scale per output channel, activations quantized per token at run time, so no calibration data. The 410 linear layers of the decoder are FP8; the embeddings, the norms and the output layer that scores the answers stay at higher precision.

Apache License 2.0: see LICENSE. The weights in this repository are modified from Apache 2.0 open weights; the change notices (NOTICE) and the training data's sources and terms (ATTRIBUTIONS.md) will be added in a later update.

@misc{salfati2026rene1,

title = {Rene-1: A 31B Decision Reader That Knows When It Is Sure},

author = {Salfati, Elon},

year = {2026},

howpublished = {\url{https://huggingface.co/salfatigroup/rene-1-31b-fp8}}

}

- Downloads last month

- -

- Balanced skill, the 37 benchmarks evaluated on Decision Index 0.2, 37 of 40 benchmarks, evaluated locallyLocal evaluation on rows rebuilt from the public sources64.180

- Balanced skill, all 40, the 3 not evaluated scored as zero on Decision Index 0.2, 37 of 40 benchmarks, evaluated locallyLocal evaluation on rows rebuilt from the public sources58.610

- Median calibration error, 10 equal-mass bins, across the 37 on Decision Index 0.2, 37 of 40 benchmarks, evaluated locallyLocal evaluation on rows rebuilt from the public sources0.046