Click, code and call tools across desktop, web and mobile.
Holo4 27B scores 85.2% on OSWorld at $0.08 per task.
Research · September 28, 2026 · 5 min read
Holo4 is our new series of agentic models. It comes in two sizes: 27B dense and 35B-A3B Mixture of Experts. Both are available on the H Models API. We are also releasing an updated version of Holotron 3: Holotron4 Nano.
Holo4 builds on our previous model and interacts with software through any available interface: GUIs, code, MCP and APIs. It scores well on academic benchmarks, but we built it for real business workflows. It was trained through supervised and reinforcement learning on a large set of environments and tasks, including those generated by our Agentic Task Factory.
Models built for every interface
Holo4 clicks and types on a screen, writes and runs its own code, and calls MCP or API tools. It uses whichever fits the task. Most agentic models are trained for one interface only: GUI-focused models are blind without a screen, while models that prefer tool calling are stuck in front of an application that has no API. Real work is not siloed that way, and a single business task can require combining these different approaches.
Holo4 runs on desktops, on the web, on Android, in a code sandbox and against business APIs. It is the same model in each case and it is called the same way. You do not need to select a different model for each platform.
Holo4 models improve significantly over their Qwen base. Holo4 trails only the strongest closed models on long workflows: on OSWorld 2.0, Holo4 27B scores 61.7% against 81.8% for Opus 5.5, and Holo4 35B-A3B reaches 30.9%. However, it does so with orders of magnitude fewer parameters and at a much lower cost. We open-source every trajectory behind our scores on public benchmarks: replay each step at trajectories.hcompany.ai or download them from Hugging Face.
Competitive with the frontier, at a fraction of the cost
On the hardest academic benchmarks for desktop control (OSWorld 2.0) and API use (AutomationBench), Holo4 competes with frontier models at a much lower cost per task.
OSWorld 2.0Average partial score (%) · higher is better
Holo
Base model
Reported
Closed frontier
← Lower costUSD per task · log scale
AI that does work
Trained on environments and tasks from our Agentic Task Factory, Holo4 models excel on professional software. The examples below show Holo4 27B alongside Qwen3.8 27B, its base model.
84 calls · 1.3M tokens
Same prompt and harness for both models.
How we built Holo4
Agentic Task Factory
Our internal set of agentic pipelines builds interactive environments and verifiable tasks from documentation alone, such as screenshots of real websites or open-source software. So far it has produced about 10,000 tasks across web apps, MCP servers and desktop environments, including hybrid environments that expose the same state through a GUI and MCP.
Training
01
Supervised fine-tuning
127B tokens. About three quarters are successful agentic trajectories from our Agentic Task Factory: desktop (45%), web (14%), MCP and API (12%) and mobile (3%). The rest covers multimodal reasoning, GUI grounding, and text-only tool use and coding.
02
Two RL experts
Asynchronous online reinforcement learning on long-horizon tasks trains two specialized LoRA experts on the fine-tuned model: one for desktop and web, one for terminal, MCP and API.
03
One merged model
Both experts merge back into the fine-tuned model with equal weight and no further training. The result combines the general skills from supervised fine-tuning with each expert's specialized skills.
Harness
Alongside training, we rebuilt our harness, the loop that executes the model's actions and manages its context over hundreds of steps, using feedback from agentic performance on OSWorld 2.0. Agents tagged why each task failed and engineers reviewed their fixes. The largest changes were giving the agent a reliable memory that can keep track of hundreds of steps, and a shell on the desktop machine itself.
Holotron4 Nano
Our post-training stack is designed to adapt to new foundation models and produce agents that generalize across interfaces and environments. As a member of the NVIDIA Nemotron Coalition, we applied our latest stack to the Nemotron 3 Nano Omni model as a follow-up to Holotron 3.
The same recipe turns Nemotron 3 Nano Omni into Holotron4 Nano, a generalist agentic model that significantly improves over the base model on GUI workflows and in environments exposing MCP, APIs or coding sandboxes.
These gains show that our recipe transfers well and can turn a generalist model into an agentic expert. Nothing in it is size-specific.
Run it yourself
Both sizes are available today on the H Models API. Weights are on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF, next to our small model, Holotron4 Nano.
We will release optimized DSpark drafter checkpoints in the coming days to further accelerate inference.
New
Flagship
Holo4 27B
Dense, 27B
Best accuracy on long, multi-step tasks across web, desktop, and mobile.