AI News
H Company released two open‑weight computer‑use agents, but their scores and licenses point builders in different directions. Here is what the public record supports.
One release, two different deployment choices
H Company released Holo4 on September 28 as a family of vision‑language models that can work through graphical interfaces, code, MCP tools and APIs. The launch includes a dense 27‑billion‑parameter checkpoint and a mixture‑of‑experts checkpoint with 35 billion total parameters and about 3 billion active at a time. Both are available through H’s hosted Models API, and downloadable versions are listed in BF16, FP8, NVFP4 and 4‑bit GGUF formats.
The practical choice is less simple than the family name suggests. H’s Holo4‑27B model card labels those weights CC BY‑NC 4.0, which permits reuse only for non‑commercial purposes under the licence’s conditions. The Holo4‑35B‑A3B card labels its weights Apache 2.0, a permissive licence that allows commercial use subject to its notice and other terms. A team choosing a checkpoint therefore has to consider the exact repository, not just whether “Holo4” is described as open weight.
Both cards list a maximum configured context of 262,144 tokens. That is a configuration limit, not evidence that either model remains reliable for every task across that entire window. The stronger evidence concerns specific agent evaluations—and those results also split the family.
The 27B checkpoint is much stronger on H’s long‑workflow run
On OSWorld 2.0, H reports a 61.7% partial‑reward score for Holo4‑27B at a mean model cost of $1.22 per task. Holo4‑35B‑A3B scores 30.9% at $0.61 per task. On those vendor‑run figures, the dense model buys 30.8 additional percentage points of partial reward for twice the reported mean model cost per attempted task.
That comparison is useful within H’s own two runs. It does not establish that either checkpoint will deliver the same completion rate or cost on a company’s workflow. OSWorld 2.0 contains 108 long, stateful professional tasks; its site says the median task takes a skilled person about 1.6 hours and distinguishes partial reward from strict binary completion. H’s public trajectory card lists 106 OSWorld 2.0 runs for each checkpoint, two fewer than the benchmark’s 108 tasks; the retained sources do not establish whether those two were excluded from scoring or only omitted from the released traces. A partial score can credit pieces of an unfinished workflow, so it should not be read as a 61.7% end‑to‑end success rate.
H’s launch post also warns that releases, harnesses and task subsets differ across the comparison chart. Its Holo4 scores came from H’s harness, while several frontier‑model points came from vendors or the official leaderboard. Independent coverage from The Register likewise cautions that a smaller or lower‑priced model is not automatically cheaper per job when long agent runs consume more reasoning tokens. The launch numbers are a reason to test the 27B model, not a substitute for measuring completed work on the target queue.
AutomationBench’s public score needs its held‑out footnote
H reports strict public‑set pass rates of 45.4% for Holo4‑27B and 34.5% for Holo4‑35B‑A3B on AutomationBench v1.0.6, at $0.05 and $0.02 per task respectively. AutomationBench counts a public task as passed only when every assertion passes; the average across the 600 scored public tasks is the public pass rate. H produced these figures with its internal harness and says 480 of the 600 tasks belong to the split from which it collected training data.
The cleaner comparison is the 120‑task subset H says it held out from that collection. There, H reports 49.3% for Holo4‑27B and 31.7% for Holo4‑35B‑A3B, compared with 40.3% for Qwen3.8‑27B and 13.1% for Qwen3.6‑35B‑A3B in the same harness. The held‑out results still favour both post‑trained models over their stated bases, but they are company‑run measurements rather than an independent replication.
The distinction between public and official scores matters. The AutomationBench repository says its public set and the private set behind the official leaderboard are separate, and that local public‑set scores may not match the private leaderboard one‑for‑one. H says it has not yet reported Holo4 on that private set. Until it does, Holo4’s numbers should not be placed directly beside private‑set leaderboard results as if they shared one test population.
The released traces improve auditability, not independence
H did publish substantially more evidence than a benchmark screenshot. Its Holo4 trajectories dataset lists 7,366 agent runs across OSWorld, OSWorld 2.0, AndroidWorld, AutomationBench, PinchBench and Agents’ Last Exam. Each trajectory record can include the task, steps, reasoning, actions, tool results, screenshots, token use and verifier outcome. The dataset is Apache 2.0, while upstream task content retains its original licence.
That lets researchers inspect failure paths, recalculate summaries and compare the two checkpoints at the run level. It also exposes a useful limitation: the dataset says credentials, internal hosts and personal data are masked, screenshots containing them are replaced, and a few tasks are omitted. Public traces make the launch more auditable; they do not turn H’s evaluation into an independent one.
A meaningful next check would rerun the downloadable checkpoint, pinned harness and benchmark environment under a separate operator, then report both partial credit and strict completion. For deployment decisions, teams should add recovery time, damaged‑state cleanup and human intervention to model‑token cost. A computer‑use agent that leaves a workflow half‑finished can be more expensive than a higher‑priced run that completes cleanly.
Which Holo4 checkpoint fits which job
For non‑commercial research and evaluation, Holo4‑27B is the stronger starting point supported by H’s published long‑workflow result. For a commercial self‑hosting review, Holo4‑35B‑A3B has the clearer permissive weight licence, but its lower OSWorld 2.0 result makes workload‑specific testing essential. Hosted API use is a separate contract question; a repository’s weight licence does not by itself define the service terms.
Before choosing either model, freeze a small test set that mirrors the intended work. Record strict completions, partial outcomes, retries, human rescue minutes, latency and total inference spend. Pin the checkpoint, quantization, harness, step limit and application versions so a later rerun measures model or system changes instead of a moving environment. Run computer‑control tests in isolated accounts with least‑privilege tools and reviewable logs.
The release is still notable: two relatively compact checkpoints can move between screens, code and structured tools, and H has made the public evaluation traces inspectable. The evidence supports a narrower conclusion than “open Holo4 beats frontier agents.” It supports testing the non‑commercial 27B model when completion quality is the priority, and testing the Apache‑licensed 35B‑A3B model when permissive self‑hosting rights are a requirement.
Figure: Holo4 checkpoint and evidence split. Sources: H Company launch post, Holo4‑27B model card and Holo4‑35B‑A3B model card, accessed September 29, 2026. Figure by OpenTools Team; no third‑party expressive material used.