On September 28, France's H Company released Holo4, a family of AI models that combine screen operation with code execution and can also call external tools. Rather than only clicking through screens designed for humans, the same model chooses how to act depending on the task. The company provides both the model weights and an API. However, on OSWorld 2.0, which evaluates long tasks, the flagship 27B model scored 61.7% when partial credit was included, and its full success rate was 41.5%. How far does the ability to handle a screen translate into the ability to finish a job? That difference also shows up when calculating the cost of choosing a cheaper model.
Holo4 connects clicking and code
Holo4-27B is a dense model with 27 billion parameters, built on Alibaba's Qwen3.8-27B. The other model, Holo4-35B-A3B, is built on Qwen3.6-35B-A3B. It is a mixture-of-experts (MoE) model that activates roughly 3 billion of its 35 billion total parameters during inference. H also released Holotron4 Nano, which is further training of NVIDIA's Nemotron 3 Nano Omni.
A GUI is a screen with buttons and input fields that you operate by looking at it. An API is an entry point for calling software functions from a program, and MCP is a common standard that lets AI use external tools. Holo4 can click and type on a screen, write and run its own code, and call functions through MCP or APIs. The company says the same model works on desktop and web as well as Android.
Consider a job that pulls information from a business system via API and enters it by hand into an older app that has no API. An AI that can only operate screens would keep clicking even where a machine-facing entry point exists. Conversely, an AI that relies only on tool calls would stall at apps that can be operated only through the screen. A model that handles both has room to keep working while changing the means available for each app. This is a use case inferred from the design, not a report that H has run this kind of work in production.
H also showed a demo of building the Eiffel Tower in FreeCAD. In software that can assemble shapes with code, combining an understanding of the interface with the ability to write programs makes sense. A successful individual example does not show that arbitrary design work can be handed off unattended.
Screen operation and local execution themselves are not new with this release. The previous-generation Holo3.1 already advertised support for web, desktop and mobile, and offered function calling and quantized weights. What stands out in Holo4 is the design: how it trains the ability to use multiple means of operation and connects it to long tasks.
Two specialized training tracks and memory for long tasks
In its training approach, H trained the ability to handle screens and the ability to use tools separately, then merged them at the end. It first performs supervised learning, then applies reinforcement learning in two separate tracks: one for desktop and web, and one for terminal commands and MCP/API use. It uses LoRA, which adjusts a model through a small set of added parameters, merges the two tracks with equal weights, and does no further training afterward, according to the company. This describes the training procedure; it does not mean the 35B model's MoE structure is split in two.
The "Agentic Task Factory," which creates training tasks, is a mechanism that generates operating environments and the jobs to be accomplished from documents and software. According to H, it has created about 10,000 tasks so far. Verifiers are required not only to fail the untouched state and pass the correctly completed state, but also to reject near-correct mistakes. Simply behaving plausibly on screen cannot confirm that a job was completed.
What turns a model's abilities into actual operation is the surrounding execution infrastructure. According to the model card, the process is a loop: pass screenshots and tool results to the model, execute the clicks or code it returns, and feed the results back. Simply loading the model weights does not make a PC start operating automatically.
H also rebuilt this infrastructure. It added memory that tracks progress over hundreds of steps, made it possible to run commands on the desktop being operated, and fixed issues such as image-processing bugs. It also expanded the number of operations and the time allowed during evaluation, to a maximum of 500 steps and 6 hours. These are upper limits, not average durations. Both making the model smarter and keeping it from losing track of intermediate state affect results on long tasks.
Cost per attempt is not the cost of finishing the job
The 27B model's OSWorld score of 85.2%, which H highlights, cannot be read as a success rate on long tasks. In a separate evaluation, OSWorld 2.0, the average score that gives credit for partial progress is reported separately from the share of tasks completed successfully. Extracting the two Holo4 models from the official evaluation table of September 28 gives the following.
| OSWorld 2.0 metric | Holo4 27B | Holo4 35B-A3B |
|---|---|---|
| Average score including partial credit | 61.7% | 30.9% |
| Full success rate | 41.5% | 12.3% |
| Model cost per attempt | $1.22 | $0.61 |
The cost is the input and output tokens counted by the company, converted at H Models API pricing. For both models, these are results of running each task once in this evaluation. The 61.7% figure does not mean "about 60% of jobs were completed." To see the share of tasks fully completed, you need to use 41.5%.
Calculating the model cost per full success, including failures, from Holo4's published figures gives about $2.94 for the 27B model and about $4.96 for the 35B-A3B model.
The calculations are 1.22 ÷ 0.415 for the 27B model and 0.61 ÷ 0.123 for the 35B model. This is equivalent to dividing the model cost of all attempts by the number of fully successful ones. The 35B model is cheaper per attempt, but using this evaluation's success rates as the denominator makes the 27B model cheaper. Because the original published values are rounded, the results are approximate.
These figures do not include the cost of running PCs or virtual environments, or the cost of people fixing failures. Nor are they the expected cost of retrying the same job repeatedly, or a prediction that real work would have the same success rate. Even so, they give a concrete answer to the question of whether cost per attempt alone is enough when choosing a model.
Rankings against other companies' models also come with conditions. H's comparison table lines up numbers measured by different companies with different execution infrastructure and inference settings, and the task sets evaluated do not match. OSWorld's own published documentation also asks that code, tasks and execution environment be aligned to the same version. A revised version 2.1 was released on September 16. Lining up numbers without checking versions does not isolate differences in the models alone.
AutomationBench, which evaluates API operation, carries a separate caveat. The results H published are for 600 public tasks, 480 of which are in the split from which the company collected training data. It gives a separate score for the remaining 120 tasks and says it will conduct the private official evaluation later. It cannot be asserted that all 480 tasks were used for training, but performance on public tasks also cannot be equated with strength on unseen real-world work.
Open-weight terms and headroom for running on-device
The 27B and 35B models have different open-weight license terms. The 27B model card lists the non-commercial CC BY-NC 4.0, while the 35B-A3B lists Apache 2.0. Even though the original Qwen is under Apache 2.0, the further-trained 27B model does not carry the same terms. For a company choosing a model to run in-house, this is an item to check alongside the performance table. H's API terms of use must be checked separately from the open-weight license.
The weights are provided in BF16 as well as FP8, NVFP4 and 4-bit GGUF formats. Quantization makes a model lighter by reducing the precision used for storage and computation, and GGUF is a format used for local inference. However, the explanation that the 35B model activates about 3 billion parameters cannot be read as meaning that only 3 billion parameters' worth of weights need to be stored.
A simple weights-only calculation at 4 bits gives about 13.5 GB for the 27B model (27 billion × 4 ÷ 8) and about 17.5 GB for the 35B model (35 billion × 4 ÷ 8). These are rough estimates from parameter counts; memory for image processing and for holding past inputs is needed separately. There is a gap between a model fitting and being able to run long tasks comfortably. It is no guarantee that real work will run on a particular 24 GB GPU.
As material for checking what happens during execution, H also published benchmark action logs. The records include the model's reasoning and actions and tool results, with screen images attached. They also let you trace whether a task succeeded, how much partial credit it received, and how much time and how many steps it used. However, personal information and the like are masked, and some images and tasks are excluded, so this is not all of the raw data published as is. The release is material for third-party verification, not itself the result of an independent reproduction test.
H also plans to release, within a few days, the weights for DSpark, which speeds up inference by using an auxiliary model. There is room for speed improvements, but when deciding whether to adopt the models, you need to line up the available means of operation and usage terms, then measure where your own work stalls and how much total cost it takes per normally completed job. Only after that verification can you judge the value of having one AI span systems that have APIs and apps that have only a screen.
