On September 17, PrismML released Ternary Bonsai 2 27B, a 27-billion-parameter model. Based on Qwen3.8-27B, it shrinks a 53.80GB language model to about 5.9GB. In the company's 20-benchmark evaluation, it kept 98.2% of the original model's score. The weights are released under Apache 2.0, and the company explicitly states support for iPhone and iPad. This widens the options for moving a 27B-class model onto everyday devices. But performance loss varies with inference settings and the type of task, and running it on a phone requires checking the distribution format and app support. Before adopting it, look beyond the average score and file size to the success rate on long tasks and the memory needed at runtime.
How It Fits into 5.9GB
Instead of reducing the number of parameters, Bonsai 2 27B cuts the number of bits used to represent each weight. The core of the approach is representing weights as ternary values (−1, 0, +1) plus an FP16 scale shared by every group of 128 weights. While preserving the structure of the original Qwen3.8-27B, this reduces both the storage size and the amount of data read from memory during inference.
Adding the shared scale to the information needed for the ternary values comes to about 1.71 bits, and including a small number of high-precision parameters, the theoretical average is 1.72 bits. Because of packing constraints in the actual distribution formats, PrismML's technical report lists the smallest format as 1.76 bits per weight and 5.93GB. The announcement's "about 5.9GB" refers to this language model.
Not everything is ternary. About 26.2 million parameters, or 0.0976% of the language model, remain in high precision, including weights used in the linear-attention state-update path and in normalization. The part that processes images is also handled separately. With these small exceptions, the whole language model is stored at low bit widths.
Another element is the Hadamard rotation, a fixed orthogonal transform. The stored weight representation is matched with a transform applied to the input side at runtime. Dedicated compute kernels use the compressed weights directly, without converting the whole model back to FP16, so the small size carries over to inference. However, the input-side transform costs computation time, and reducing it remains a development task.
Against the previous generation, the update of the base model also contributes. PrismML says the performance retention rose from about 95% to 98.2%, but it has not separated how much came from progress in the base model and how much from improvements in the compression method. The generation-to-generation figures cannot be read as an improvement from compression technology alone.
Separating Where the 98.2% Applies
In the 20-benchmark evaluation, Bonsai 2 27B averaged 83.9, compared with 85.4 for the Qwen3.8-27B FP16 baseline. The measurements were made by PrismML itself and span math and code generation through image understanding. The reasoning effort used was xhigh; results for medium, which shortens reasoning, are also in the technical report.
Calculating retention from PrismML's published figures gives 98.2% for the 20-benchmark average at xhigh, 96.0% at medium, 75.8% for Terminal-Bench 2.1, and 75.4% for SWE-bench Verified.
| Benchmark and setting | Qwen3.8-27B FP16 | Bonsai 2 27B | Retention vs. original |
|---|---|---|---|
| 20-benchmark average, xhigh | 85.4 | 83.9 | 98.2% |
| 20-benchmark average, medium | 82.6 | 79.3 | 96.0% |
| Terminal-Bench 2.1, separate agent evaluation | 69.7 | 52.8 | 75.8% |
| SWE-bench Verified, separate agent evaluation | 80.6 | 60.8 | 75.4% |
Source: Section 4 and Appendices B and C of PrismML's September 2026 technical report. Retention is each row's "Bonsai's published score ÷ FP16's published score × 100," rounded to one decimal place. The bottom two evaluations are not included in the 20-benchmark average, and scores from different evaluations have not been combined.
Even at medium, the average scores are close, but the 98.2% at xhigh cannot be applied as is. If you shorten reasoning to cut waiting time in a local environment, the trade-off with quality changes. Moreover, the evaluation allows an output budget of up to 81,920 tokens for some tasks. That does not mean the model always produces output that long, but it is also different from an evaluation of short answers only.
The gap is larger in agent evaluations involving long tasks. Terminal-Bench tests solving tasks by operating a terminal, and SWE-bench Verified tests fixing software bugs. PrismML says it evaluated all 89 Terminal-Bench tasks with one attempt each using Terminus-2, and all 500 SWE-bench items using mini-swe-agent. Good single-shot code generation or function calling does not guarantee a similar rate of completing a task from start to finish.
Still, it is valuable that a model of about 6GB can complete such tasks in some cases. A high average retention rate and being a drop-in replacement for the original model should be judged separately. The figures in the table are the developer's own measurements, not independent tests confirming the same success rates or run times on ordinary PCs.
Also, the Hugging Face GGUF model card, as checked on September 20, listed a different aggregate: 84.78 versus 86.32 across 14 benchmarks. The ratio is also 98.2%, but the set of benchmarks differs. Individual values on the model card should not be spliced into the 20-benchmark aggregate from the announcement and technical report.
The Smallest Format Isn't Always the Fastest
The distributed weights include GGUF PTQ1_0, which prioritizes size, and PQ2_0, which places ternary values in 2-bit slots to make them easier to extract. The Apple-oriented MLX version stores the same ternary values in a different format. Checking the model cards and distributed file sizes as of September 20, the sizes are as follows.
| Format / configuration | Storage size | What is included |
|---|---|---|
| GGUF PTQ1_0 | About 5.95GB | Language model |
| GGUF PQ2_0 | About 7.21GB | Language model |
| GGUF extra image file | About 0.63GB | Processing component used for image input |
| MLX 2-bit | About 8.60GB | Language model and image-processing component |
Sources: the GGUF distribution page and the MLX distribution page. Sizes are in decimal GB and are not the total memory required at runtime. The technical report lists 5.93GB/7.25GB for GGUF and 8.49GB for MLX; since these differ from the distribution side, the table uses the distribution-side values as of the check date.
The MLX version stores a scale and a bias for every group of 128 weights, so it takes more storage than the information needed for the ternary representation alone. It also bundles the image-processing component. Assuming from the "5.9GB model" headline that every distribution format fits in the same size would lead to a wrong estimate when deploying.
Speed also depends on the format. In the same-version measurements in the technical report, generation speed on an RTX 5090 was 142.5 tokens/second for PQ2_0 and 134.4 tokens/second for PTQ1_0. On an RTX 4090, however, the figures were 90.9 and 96.7 tokens/second respectively, so the smaller format wins. This is because the benefit of transferring fewer weights and the computational burden of unpacking the packed ternary values differ by GPU.
Speed at reading long prompts also deserves attention. On the RTX 5090, input processing ran at 4,121 tokens/second with PQ2_0 and 1,901 tokens/second with PTQ1_0. Choosing the format that generates faster does not necessarily shorten the wait for the first answer when you hand over a long document.
These values were measured on September 16, 2026, with batch size 1, no accumulated context, and no image-processing component. Generation used 128 tokens and input processing used 512 tokens, averaged over three measurements after a warm-up. The M5 Max figure of 46.8 tokens/second is also a result of running PQ2_0 on llama.cpp's Metal backend, not a measurement of the MLX version. None of these numbers guarantees speed at maximum context length or the time until an answer is complete.
The power-efficiency comparison also has a measurement scope. The technical report's RTX 4090 PQ2_0 figure of 0.714mWh per token is for the entire board including the GPU and its memory, not the whole PC's power consumption. Apple-side measurements exclude DRAM, so the report does not directly compare per-token power between the two.
It Runs on iPhone, but What to Check
In this announcement, PrismML states explicitly that Bonsai 2 27B runs on Mac, iPhone, and iPad through MLX and its own low-bit compute kernels. Smartphones are among the destinations for bringing a large model onto local devices.
However, the concrete iPhone measurements need to be read with the generation in mind. The "11 tokens per second on iPhone 17 Pro" the company gave in its July 14 announcement was for the previous-generation "1-bit Bonsai 27B," at 3.9GB. It is not the speed of the ternary Bonsai 2 27B this time. The current announcement and technical report do not give specific supported iPhone models, on-device generation speed, or peak runtime memory use.
The way to look at size also changes. The announcement's roughly 5.9GB refers to the language model in the smallest GGUF format, while the MLX package listed above is about 8.60GB. Working memory is needed beyond the stored weights, and the OS and apps are running on the same device. So you cannot conclude that an iPhone with 5.9GB of free space can use it. The headroom required also changes depending on whether you handle images and how long a text you feed it.
As an app for trying it, PrismML's official guide introduces "Locally AI." The App Store release notes also mention adding Bonsai 27B for supported devices. But this should be considered separately from confirming support for Bonsai 2 27B. Choosing without checking the model generation means trying something different from the performance discussed in this article.
Caveats remain in the developer documentation, too. The official Bonsai 2 documentation says it runs on standard MLX, but the runtime note inside the distribution says the standard loading path does not apply the required conversion, and that loading the whole model in Swift needs additional integration. Given the discrepancy in the explanations, the "MLX-compatible" label alone does not guarantee it will run immediately when loaded into an ordinary iPhone app.
Even so, there is real significance in keeping a model of this scale on a phone. For example, when summarizing documents on the device or writing text based on their contents, it can reduce the need to send the full input to a cloud inference service. With a setup that downloads the model and completes processing on the device, designs that work in places without an internet connection also become possible. Communication for features that call external search or cloud APIs is separate, but being able to choose where to process personal documents is a use value distinct from the size reductions for PCs.
To judge practicality, check not only whether it launches but also the wait for the first response and the heat and battery drain when generating long answers. Short generation tests on PCs and the previous generation's iPhone measurements cannot answer that. With the iPhone support announcement as a starting point, the next evaluation is which jobs can be handed over comfortably on the model and app you use.
Adoption Depends on the Task and Context Length
To run Bonsai 2's GGUF, PrismML's build of llama.cpp is currently required. According to the official run guide, the standard build does not yet include the needed Hadamard transform processing, and PQ2_0 and PTQ1_0 are refused at load time. The GGUF extension alone does not mean an app you already use is compatible.
The maximum context length of 262,144 tokens also does not mean it can always be used with little memory. Beyond the weights, a KV cache for referencing past input and working space are needed. The official guide puts the FP16 KV cache for the 27B family at 64KiB per token, so additional capacity grows as the context lengthens. This is why the launch script chooses the context length according to installed memory.
When a language model shrinks to about 5.9GB, options open up for PCs that previously struggled even to hold the weights. But the required configuration changes depending on whether you spend the freed memory on long context or on image input. Folding storage size, generation speed, and task success rate into a single "efficiency" figure obscures that choice.
If you try it, check accuracy and time to a complete answer at each reasoning effort, using your own documents and code. If you plan to hand it autonomous bug fixing, measure the rate at which it completes the task rather than single answers. If it gives sufficient results under those conditions, Bonsai 2 27B is a strong candidate for moving work you want to process without sending it to the cloud onto your own devices.
