Bonsai 2 27B on paperwork-v3: 0/9 resolved without thinking, 6/9 with (72.2% practical)
Bonsai 2 27B is PrismML's ternary successor on a Qwen3.8-27B backbone. On paperwork-v3 it scored 0/9 strict without thinking and 6/9 strict, 7/9 core (72.2% practical) at medium thinking effort — a 44.4-point mode split.

Bonsai 2 27B is PrismML's ternary-weight successor to Bonsai 27B, built on a Qwen3.8-27B hybrid-attention backbone. On the nine-case paperwork-v3 suite it produced two very different scoreboards: with reasoning switched off it closed 0/9 strictly and 5/9 core-oracle for 27.8% practical; with thinking at the vendor-recommended medium effort it closed 6/9 strictly and 7/9 core-oracle for 72.2%. The 44.4-point delta is the story.
Bonsai 2 27B is the second-generation ternary model from PrismML. Where Bonsai 27B compressed a Qwen3.6-27B checkpoint, Bonsai 2 keeps ternary weights across the transformer matrices of a Qwen3.8-27B hybrid-attention backbone and ships as a ~7.2 GB PQ2_0 GGUF that needs PrismML's own llama.cpp fork to run. The vendor claims 98.2% of FP16 intelligence retained — a figure from its 20-benchmark suite (83.9 vs 85.4 for the FP16 backbone); the GGUF card quotes an 84.78 average across 14 thinking-mode benchmarks. We ran it on the same Mac mini M4 with 64 GB that hosts the rest of the local rows, against the same nine paperwork-v3 cases, twice: once with reasoning disabled, once with thinking at medium effort.
Two runs, two scoreboards
Local Model Bench separates two signals per case: a strict resolved signal (the model produced a final artifact the hidden oracle accepted as exact) and a core-oracle pass signal (the model identified the right audit facts and only failed on cosmetic or arithmetic closure). Practical score = 0.5 × (strict/9) + 0.5 × (core/9).
| Mode | Strict resolved | Core-oracle | Practical |
|---|---|---|---|
| No-think (benchmark convention) | 0/9 | 5/9 | 27.8% |
| Thinking, medium effort | 6/9 | 7/9 | 72.2% |
Every other model on the leaderboard ran under the no-think convention: the harness sends think: false and enable_thinking: false per request and the prompt ends with an explicit "do not think out loud" instruction. Bonsai 2 is the first model in the suite where flipping that single switch changes the result by 44.4 points.
Update 2026-09-25: a later audit found that several leaderboard rows, including Qwen3.6 27B, Gemma-4 26B and Muse-Glimmer, produced hidden reasoning tokens even though these flags were sent — the flags are ignored by their chat templates. Bonsai 2's no-think run was verified clean (zero reasoning tokens), so the 44.4-point comparison above stands. See The no-think leaderboard that wasn't.
What no-think failure looks like
Without a thinking trace, three of the five generated-invoice cases (P01, P02, P05) never produced a parseable audit at all. The model's document reading was correct. The right invoices approved, the right ones flagged for review. But the final proof_code field came back as a raw arithmetic expression instead of a number:
"proof_code": 18737 + 7801 + 7802 + 8422 + 97 * 2The expression was even correct. It evaluates to 42,956, which is the oracle value. The model constructed the right formula and never computed it. Invalid JSON means no audit artifact, which means zero checks. The two remaining generated cases returned valid JSON with wrong arithmetic (proof_code 100000 and 100), and all four workflow cases reached core-oracle pass but failed strict on the same proof-code and manifest steps.
What thinking changes
At medium reasoning effort the same five generated-invoice cases went 4/5 strict. P01, P02, P04 and P05 closed cleanly, proof codes evaluated and exact. P03 failed on a different signature: invoice misclassification plus a missed duplicate-risk warning, which is a document-reading failure, not an arithmetic one.
The workflow cases moved less uniformly. W06 remittance-split and W07 credit-offset closed strictly with all four oracle checks green. W04 messy-intake stayed at core-oracle pass on a manifest error. W05 email-attachment intake actually got worse than its no-think run: visible checks passed, but the final document set, an ignored-document ID, the proof code and the proof file all missed. On nine cases with one run each, single-case regressions like that sit inside run-to-run noise.
Side by side
| Model | Strict | Core | Practical | Mode |
|---|---|---|---|---|
| Bonsai 2 27B (thinking) | 6/9 | 7/9 | 72.2% | local, ternary PQ2_0 |
| Qwen3.6 27B | 5/9 | 8/9 | 72.2% | local, FP16-class |
| Qwen3.8 27B | 3/9 | 7/9 | 55.6% | local/cloud, same backbone |
| Bonsai 27B (v1) | 2/9 | 4/9 | 33.3% | local |
| Bonsai 2 27B (no-think) | 0/9 | 5/9 | 27.8% | local, reasoning off |
Read strictly on the numbers, Bonsai 2 with thinking ties Qwen3.6 27B at the top of the local rows and beats its own Qwen3.8-27B backbone by 16.6 points. Read conservatively, nine cases with one run per case cannot separate a 6/9 from a 5/9. What the table does show cleanly is the mode split, and that is a measurement with zero ambiguity: same model, same cases, same day, same harness, one flag flipped.
Why the split is so large
The proof-code formula is deterministic: sum of approved gross cents, plus the numeric parts of every invoice ID across the verdict lists, plus 97 times the warning count. It is multi-digit addition with five to eight terms. A model that writes the formula correctly but must emit the evaluated result in the same forward pass is doing mental arithmetic under a hard constraint; a model allowed a thinking trace can compute the sum step by step and copy the result. Bonsai 2's ternary arithmetic path appears to lean on that trace harder than the dense baselines did. Its predecessor at 33.3% failed on classification and evidence shape, not on JSON validity.
The vendor knows this: the model ships with xhigh reasoning effort as the template default, supports medium as the recommended balance, and does not support low at all. Our no-think run is therefore a configuration the model was not designed for. It is still a legitimate measurement, since the benchmark convention exists to compare what a user sees without reasoning traces and several production deployments disable thinking for latency. But the fairest reading of Bonsai 2 is the thinking row.
What this means for the leaderboard
We record both rows. The no-think row (27.8%) keeps Bonsai 2 comparable with every model that ran under the benchmark convention. The thinking row (72.2%) records what the model does with thinking on, at the documented medium balance point. If a leaderboard needs one number per model, the thinking number is closer to real use; if it needs one rule for all models, the no-think number is the honest one.
What we would still want to know
A run at xhigh effort would show whether medium already saturates the suite or leaves headroom. A second thinking run would separate the W05 regression from noise. And a re-run of Qwen3.8 27B with thinking enabled would answer the obvious follow-up: is the 44-point delta a Bonsai-2 property or a property of the shared backbone that dense Qwen3.8 also carries.
Where it worked
- 4/5 generated-invoice cases closed strictly with thinking, including the three that produced invalid JSON without it.
- W06 and W07 closed strictly, the first fully clean workflow runs for any ternary model on this suite.
- All four workflow cases reached at least visible-checks pass in both modes; no run produced a zero.
- ~16.5 tok/s generation on Mac mini M4 (PQ2_0, Metal), ~80–160 s per generated case no-think.
Where it failed
- 0/5 generated cases produced a valid audit artifact without thinking; three returned an unevaluated arithmetic expression as proof_code.
- P03 failed in both modes, the only case where thinking changed the failure signature from arithmetic to misclassification.
- W05 thinking run regressed versus no-think: final document set, ignored-document ID, proof code and proof file all missed.
- The GGUF requires the PrismML llama.cpp fork; stock llama.cpp rejects the PQ2_0/PTQ1_0 packings, and the vision tower ships as a separate mmproj file.
How the model is positioned
- PrismML positions Bonsai 2 as "full 27B-class reasoning in ternary transformer weights" — a vendor claim, not a benchmark result.
- The GGUF needs the vendor's llama.cpp fork (CUDA + Metal); the MLX-2bit companion build runs on stock MLX. Stock GGUF loaders, including LM Studio's and Ollama's, do not load these files.
- The template defaults to xhigh reasoning effort; medium is the documented balance point; low is unsupported.
- Apache-2.0, 262K context, vision via a separate mmproj pack (~0.63 GB in a Q8_0 container; a BF16 reference is also shipped).
- Local Model Bench tests one specific practical workload, not the full model card.
What was actually tested
- All five generated-invoice Paperwork Trial cases plus all four agentic Paperwork Workflow cases, twice: once with reasoning disabled server-side, once at medium reasoning effort.
- Runtime: PrismML llama.cpp fork b10735 (macOS arm64 prebuilt), PQ2_0 packing (7.2 GB) plus Q8_0 mmproj vision pack, full GPU offload, 64K context, temperature 0 / top_p 1 per benchmark convention.
- Score check: 0.5×0/9 + 0.5×5/9 = 27.8%; 0.5×6/9 + 0.5×7/9 = 72.2%.
- Comparisons are the same nine cases on the same Mac mini M4 hardware class with the same suite version.
Verdict
Bonsai 2 27B on paperwork-v3 is two benchmark results in one model. Without a thinking trace it cannot close the arithmetic last mile and lands at 27.8% practical, behind its own predecessor. At the vendor's recommended medium thinking effort it closes 6/9 strictly for 72.2% — level with Qwen3.6 27B and ahead of the Qwen3.8-27B backbone it is built on, in a 7.2 GB package that generates ~16.5 tokens per second on a Mac mini. The honest summary is that ternary weights do not lose the documents; without thinking they lose the last addition. Which of the two rows belongs on a leaderboard depends on which convention the leaderboard applies.