Click, code and call tools across desktop, web and mobile.

Holo4 27B scores 85.2% on OSWorld at $0.08 per task.

Research · September 28, 2026 · 5 min read

Holo4 is our new series of agentic models. It comes in two sizes: 27B dense and 35B-A3B Mixture of Experts. Both are available on the H Models API. We are also releasing an updated version of Holotron 3: Holotron4 Nano.

Holo4 builds on our previous model and interacts with software through any available interface: GUIs, code, MCP and APIs. It scores well on academic benchmarks, but we built it for real business workflows. It was trained through supervised and reinforcement learning on a large set of environments and tasks, including those generated by our Agentic Task Factory.

Models built for every interface

Holo4 clicks and types on a screen, writes and runs its own code, and calls MCP or API tools. It uses whichever fits the task. Most agentic models are trained for one interface only: GUI-focused models are blind without a screen, while models that prefer tool calling are stuck in front of an application that has no API. Real work is not siloed that way, and a single business task can require combining these different approaches.

Holo4 runs on desktops, on the web, on Android, in a code sandbox and against business APIs. It is the same model in each case and it is called the same way. You do not need to select a different model for each platform.

Holo4 models improve significantly over their Qwen base. Holo4 trails only the strongest closed models on long workflows: on OSWorld 2.0, Holo4 27B scores 61.7% against 81.8% for Opus 5.5, and Holo4 35B-A3B reaches 30.9%. However, it does so with orders of magnitude fewer parameters and at a much lower cost. We open-source every trajectory behind our scores on public benchmarks: replay each step at trajectories.hcompany.ai or download them from Hugging Face.

Competitive with the frontier, at a fraction of the cost

On the hardest academic benchmarks for desktop control (OSWorld 2.0) and API use (AutomationBench), Holo4 competes with frontier models at a much lower cost per task.

OSWorld 2.0Average partial score (%) · higher is better

Holo

Base model

Reported

Closed frontier

← Lower costUSD per task · log scale

AI that does work

Trained on environments and tasks from our Agentic Task Factory, Holo4 models excel on professional software. The examples below show Holo4 27B alongside Qwen3.8 27B, its base model.

84 calls · 1.3M tokens

Same prompt and harness for both models.

How we built Holo4

Agentic Task Factory

Our internal set of agentic pipelines builds interactive environments and verifiable tasks from documentation alone, such as screenshots of real websites or open-source software. So far it has produced about 10,000 tasks across web apps, MCP servers and desktop environments, including hybrid environments that expose the same state through a GUI and MCP.

Training

01

Supervised fine-tuning

127B tokens. About three quarters are successful agentic trajectories from our Agentic Task Factory: desktop (45%), web (14%), MCP and API (12%) and mobile (3%). The rest covers multimodal reasoning, GUI grounding, and text-only tool use and coding.

02

Two RL experts

Asynchronous online reinforcement learning on long-horizon tasks trains two specialized LoRA experts on the fine-tuned model: one for desktop and web, one for terminal, MCP and API.

03

One merged model

Both experts merge back into the fine-tuned model with equal weight and no further training. The result combines the general skills from supervised fine-tuning with each expert's specialized skills.

Harness

Alongside training, we rebuilt our harness, the loop that executes the model's actions and manages its context over hundreds of steps, using feedback from agentic performance on OSWorld 2.0. Agents tagged why each task failed and engineers reviewed their fixes. The largest changes were giving the agent a reliable memory that can keep track of hundreds of steps, and a shell on the desktop machine itself.

Holotron4 Nano

Our post-training stack is designed to adapt to new foundation models and produce agents that generalize across interfaces and environments. As a member of the NVIDIA Nemotron Coalition, we applied our latest stack to the Nemotron 3 Nano Omni model as a follow-up to Holotron 3.

The same recipe turns Nemotron 3 Nano Omni into Holotron4 Nano, a generalist agentic model that significantly improves over the base model on GUI workflows and in environments exposing MCP, APIs or coding sandboxes.

These gains show that our recipe transfers well and can turn a generalist model into an agentic expert. Nothing in it is size-specific.

Run it yourself

Both sizes are available today on the H Models API. Weights are on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF, next to our small model, Holotron4 Nano.

We will release optimized DSpark drafter checkpoints in the coming days to further accelerate inference.

New

Flagship

Holo4 27B

Dense, 27B

Best accuracy on long, multi-step tasks across web, desktop, and mobile.