OrcaSAQ2 27B

High-fidelity 3-bit Qwen3.8 for long-horizon agents.

54 GB → 12.3 GB · +0.02% PPL · 93.2% Top-1 Agreement · 0.031 KLD · 262K Context

OrcaRouter AI Gateway · X · Discord · GitHub · All Models

OrcaSAQ2 27B compresses Qwen3.8-27B from a 54 GB BF16 checkpoint to 12.3 GB while preserving extremely high fidelity to the original model.

Built for: long-horizon agents · coding · tool use · reasoning · stateful execution

OrcaSAQ2 is a proprietary sensitivity-aware mixed-precision quantization system developed by OrcaRouter and its research team behind.

It is optimized around one goal: Preserve as much useful model behavior as possible inside a practical GPU memory envelope.

The resulting checkpoint provides:

- 77.2% smaller storage footprint

- only +0.02% perplexity versus BF16

- 93.2% token-level Top-1 agreement

- 0.031 mean KLD

- 262K context

- thinking mode

- tool calling

- MTP speculative decoding

- production serving through vLLM

4.4× smaller. +0.02% perplexity.

The point is not 3-bit.

The point is what survives at 3-bit.

All numbers below are measured using these exact OrcaSAQ2 weights against the BF16 reference through the same evaluation path.

16,376 predicted tokens

BF16 5.6468 ████████████████████████████████████████

OrcaSAQ2 5.6482 ████████████████████████████████████████

Delta: +0.02%

OrcaSAQ2 vs BF16

███████████████████████████████████████████████░░░ 93.2%

Qwen3.8-27B BF16

██████████████████████████████████████████████████ 54.0 GB

OrcaSAQ2

███████████ 12.3 GB

77.2% smaller.

Short benchmarks can hide small degradation.

Agents cannot.

A small model error can change a tool call.

That changes the environment state.

The changed state affects every decision that follows.

Plan

↓

Act

↓

Observe

↓

Decide

↓

Recover

↓

Repeat

↓

...

↓

Task Success

Across long trajectories, small errors can compound into large behavioral differences.

That makes long-horizon execution an especially useful stress test for compressed reasoning models.

OrcaSAQ2 performs strongly on long-horizon workloads relative to models in its deployment and parameter class, despite operating from a 12.3 GB checkpoint.

This makes it particularly suitable for:

- coding agents

- terminal agents

- browser agents

- computer-use agents

- security agents

- repository-scale tasks

- multi-tool workflows

- failure recovery

- long-running stateful execution

Perplexity asks:

How similar is the next-token distribution?

Long-horizon evaluation asks:

Can the model still finish the job after many decisions?

For agent models, both matter.

Agent benchmarks depend heavily on the surrounding scaffold, tools, reasoning budget, timeouts and execution environment. The results below are therefore shown as public reference points, not direct apples-to-apples comparisons.

70.0% SWE-bench Verified from a 12.06 GB 27B checkpoint.

58.4% Terminal-Bench 2.1 while fitting in ~12 GB of checkpoint storage.

Public scores use different agent stacks and should not be interpreted as a strict model-only ranking.

Measured under a 15.7 GiB GPU memory cap.

Single-stream decode

MTP off █████████████████████████████ 65.3 tok/s

MTP on ████████████████████████████████████████

90.1 tok/s

+38% single-stream decode throughput

MTP trades additional compute and KV capacity for stronger interactive decode performance.

It is particularly useful for:

- coding assistants

- interactive agents

- terminal agents

- tool-heavy applications

- low-concurrency inference

For highly batched workloads, benchmark both configurations.

OrcaSAQ2's checkpoint is 12.3 GB.

That makes deployment possible on hardware that cannot hold the original 54 GB BF16 checkpoint.

16 GB GPU

┌───────────────────────────────────────────┐

│ │

│ OrcaSAQ2 weights 12.3 GB │

│ ███████████████████████████████████ │

│ │

│ Remaining ~3.7 GB │

│ ██████████ │

│ │

└───────────────────────────────────────────┘

Actual usable memory depends on:

- vLLM overhead

- KV-cache configuration

- MTP

- batch size

- context length

- CUDA graph configuration

A practical starting point for a 16 GB GPU is approximately 32K interactive context, then tune based on the workload.

The model architecture supports up to 262K context.

plan → act → observe → recover → repeat

Repository-scale generation, editing, testing and debugging.

Structured workflows where action-selection quality matters.

Preserving the capabilities of the 27B base model under an aggressive deployment constraint.

A 12.3 GB checkpoint designed around practical inference hardware.

vLLM + MTP + OpenAI-compatible APIs.

One prompt each, first attempt.

The standard SVG test, asked for as an animation.

Chain over the chainring, cranks 180° out of phase, parallax background. Pure SMIL, no JavaScript. Used as generated.

Create a html low-poly 3D models of the Statue of Liberty

A single self-contained HTML file: Three.js scene, orbit controls, procedural geometry.

pip install -U vllm huggingface_hub

pip install git+https://github.com/Continuum-AI-Corp/OrcaSAQ2-kernel

hf download orcarouter/OrcaSAQ2-27B \

--local-dir ./OrcaSAQ2-27B

vllm serve ./OrcaSAQ2-27B \

--served-model-name OrcaSAQ2-27B \

--max-model-len 262144 \

--reasoning-parser qwen3 \

--enable-auto-tool-choice \

--tool-call-parser qwen3_xml \

--speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":2}'

from openai import OpenAI

client = OpenAI(

base_url="http://localhost:8000/v1",

api_key="not-needed",

)

response = client.chat.completions.create(

model="OrcaSAQ2-27B",

messages=[

{

"role": "user",

"content": "Analyze this repository and plan the next five actions."

}

],

)

print(response.choices[0].message.content)

temperature = 1.0

top_p = 0.95

top_k = 20

Thinking mode is enabled by default.

For agent deployments, benchmark against the actual tool schema, context distribution and reasoning budget used in production.

A low-bit reasoning model should not be judged by checkpoint size alone.

We look at the intersection of:

Footprint × BF16 Fidelity × Capability × Long-Horizon Stability × Serving Performance

A useful low-bit model must remain useful after compression.

Perplexity is useful and reproducible.

It is not a complete measure of agentic capability.

Quantization can affect:

reasoning

↓

planning

↓

tool selection

↓

state tracking

↓

recovery

↓

task completion

That is why OrcaSAQ2 reports BF16 fidelity metrics alongside downstream and long-horizon evaluation.

OrcaSAQ2 uses a proprietary sensitivity-aware mixed-precision quantization system developed by OrcaRouter.

The implementation is optimized to preserve model quality under a strict deployment-memory target.

Detailed quantization methodology, calibration strategy, precision allocation and packing techniques are not currently disclosed.

- OrcaSAQ2 inherits the capabilities, biases and limitations of Qwen3.8-27B.

- Quantization is not mathematically lossless.

- 93.2% Top-1 agreement means some token decisions differ from BF16.

- +0.02% PPL is a model-fidelity measurement and does not guarantee identical downstream performance.

- Long-horizon comparisons should use a controlled same-harness evaluation.

- This checkpoint is text-only.

- The vision tower is not included.

- OrcaSAQ2 requires the OrcaSAQ2 vLLM integration.

- Maximum architectural context does not imply that the full context fits into every GPU memory envelope.

Open multi-model code review.

Record, replay, fork and debug AI-agent runs.

Self-hosted multi-model AI infrastructure.

Open model. Open harness. Open bill.

@misc{qwen38,

title = {Qwen3.8-Max: A New Bar for Coding and Cowork},

author = {{Qwen Team}},

year = {2026},

month = {August},

url = {https://qwen.ai/blog?id=qwen3.8}

}

Apache-2.0

Inherited from:

Quantization does not change the underlying license obligations.

One Gateway. Every Model.

Route Smarter · Ship Safer · Spend Less

- Downloads last month

- 1,663