H Company launches Holo4: one model that drives GUI, code and API without swapping weights

The model that refuses to pick a lane
H Company is releasing Holo4, a family of agentic models designed to end the forced split between interfaces. Instead of making developers choose between a model that can click a screen and one that only speaks API, Holo4 does both — and writes and executes its own code, and calls tools via MCP. Same weights, same call, whether it runs on desktop, web, Android, a code sandbox or a business API. This isn't cosmetic; it's an architectural shift that lets the model flow through real tasks that don't respect artificial boundaries.
Two sizes, same approach
The family comes in two configurations: Holo4-27B dense and Holo4-35B-A3B, a mixture-of-experts architecture with 35 billion total parameters and 3 billion active per forward pass. Both are available through the H Models API, alongside an updated version of the smaller model, Holotron4 Nano. Weights are open (open weights) in FP16, FP8 and GGUF formats, and the company is publishing full traces behind its public benchmark scores — every step reproducible in a dedicated viewer or downloaded from Hugging Face. It isn't fully open source (no training code), but the transparency exceeds what most closed labs offer.
Benchmarks: chasing the giants with fewer parameters
On OSWorld 2.0, currently the toughest desktop-control benchmark, Holo4-27B scores 61.7% success versus 81.8% for Anthropic's closed Opus 5.5. The smaller MoE variant, 35B-A3B, reaches 30.9%. The gap exists, but H Company stresses it achieves these results with orders of magnitude fewer parameters and significantly lower cost per task. On cost-performance charts the company positions Holo4 as a non-dominated point against closed models — meaning no closed model is both cheaper and more accurate at the same time. On AutomationBench, which tests API usage, the picture is similar: competitive at the frontier, cheaper per task. Competitor numbers for open models (Qwen 3.8 27B and Qwen 3.6 35B-A3B) were measured in H Company's own harness; closed-model numbers come from official leaderboards, so direct comparison warrants caution, but the direction is clear.
Training in Agentic Task Factory: real environments, not toys
What separates Holo4 from its base (Qwen) isn't just fine-tuning on static data. The models were trained with supervised learning and reinforcement learning on a large set of environments and tasks generated by the company's Agentic Task Factory, an engine that produces diverse scenarios in professional software such as FreeCAD, not just toy tasks in academic benchmarks. The example the company highlights: building a 3D model of the Eiffel Tower in FreeCAD at 1 mm per meter scale, centered and axis-aligned, with a variable width profile along the height and four identical legs — a task requiring spatial reasoning, engineering precision and a long action sequence. Holo4-27B succeeds; the base Qwen 3.8 27B fails. That is the strongest signal yet that the agentic training works, not just the pre-training.
What this means in practice for developers
The practical upshot is simple: teams can stop maintaining a chain of specialized models per platform. The same Holo4 that clicks a button in a web app can, in the same run, write a Python script, call an MCP tool, then hit an internal API — no model swap, no prompt rewrite, no complex orchestration. The company is opening full traces for replay, which means learning from the model's failures as much as its successes. For anyone building production-ready automations, that removes an entire middleware layer. The only open question left