Bonsai 27B on paperwork-v3: 2/9 resolved, 4/9 core (33.3% practical)
prism-ml/bonsai-27b scored 2/9 strict and 4/9 core on paperwork-v3 (33.3% practical), ranking #10 on the leaderboard. Like Muse-Glimmer, it required an 8,000 max_tokens ceiling, but its llama.cpp template ignored reasoning flags. Failures were shape and evidence errors, not token stalls.
prism-ml/bonsai-27b is a 27B reasoning model loaded into LM Studio and called through the OpenAI-compatible endpoint at http://localhost:1234/v1 with non-default sampling settings. On the nine-case paperwork-v3 suite it scored 2/9 strict and 4/9 core, for a practical_score of 33.3%, ranking tenth on the public leaderboard. Same harness, same enable_thinking: false flag, same shape of initial failure as the Muse-Glimmer run — except the llama.cpp template bundled with Bonsai reads the flag and ignores it.
What ran
prism-ml/bonsai-27b was loaded in LM Studio and called through the OpenAI-compatible endpoint at http://localhost:1234/v1 with non-default sampling settings (temperature: 0, top_p: 1). Five generated-invoice cases (p01–p05) and four paperwork-workflow cases (w04–w07) ran sequentially against the paperwork-v3 bench, mirroring the harness configuration used for the Muse-Glimmer run.
| Case | Type | core_ok | hidden_oracle | elapsed | completion tokens | of which reasoning |
|---|---|---|---|---|---|---|
| p01 (generated_invoice_case_01) | image | yes | yes | ~6:00 | 5 563 | 5 307 |
| p02 (generated_invoice_case_02) | image | yes | yes | ~5:30 | 5 234 | 5 005 |
| p03 (generated_invoice_case_03) | image | no | no | ~9:00 | 7 801 | 7 386 |
| p04 (generated_invoice_case_04) | image | yes | no (path format) | ~6:30 | 5 889 | 5 612 |
| p05 (generated_invoice_case_05) | image | no | no | ~5:30 | 5 122 | 4 937 |
| w04 (messy_intake_workflow_case_04) | workflow | no | no | ~8:00 | n/a | n/a |
| w05 (email_attachment_intake_case_05) | workflow | no | no | ~10:00 | n/a | n/a |
| w06 (remittance_split_case_06) | workflow | yes | no (evidence) | ~9:00 | n/a | n/a |
| w07 (credit_offset_case_07) | workflow | no (no output) | no | timeout | n/a | n/a |
Two clean hidden-oracle passes (p01, p02), two core-only passes (p04, w06) where audit fields matched ground truth but an artifact shape or evidence list failed its hidden-oracle structural check, and five core fails. That counts as 2/9 strict (22.2%) and 4/9 core (44.4%), for an aggregate practical_score of 0.5 × (2/9) + 0.5 × (4/9) = 6/18 = 33.3%.
Bonsai's template ignored the reasoning-suppression flag
The Muse-Glimmer run established that LM Studio reasoning models in this harness need either a larger max_tokens ceiling or a working chat_template_kwargs: {"enable_thinking": false} flag. Muse-Glimmer's llama.cpp template respected the flag and stopped spending tokens on the hidden reasoning trace. Bonsai's template reads the request field and does nothing with it.
The diagnostic: a single Bonsai run with max_tokens: 2400 and chat_template_kwargs: {"enable_thinking": false} produced the same shape of failure the Muse run did on the first attempt — finish_reason: length, content: "", the reasoning trace eating the budget. Bumping max_tokens to 8 000 produced a clean JSON body. Sending the template-kwargs flag on top did not change the trace length.
| max_tokens | template_kwargs | reasoning_tokens | completion_tokens | content | finish_reason | core_ok |
|---|---|---|---|---|---|---|
| 2 400 | {"enable_thinking": false} | ~2 400 | 2 400 | "" | length | no |
| 8 000 | {"enable_thinking": false} | 5 307 | 5 563 | full JSON | stop | yes |
Several reasoning-suppression flags — the top-level chat_template_kwargs variant and the OpenAI-API equivalents — were all tested on Bonsai. None of them reduced the trace length. The only knob that did was max_tokens. Same shape as Muse-Glimmer's first run, but no alternative intervention to recover.
Reasoning-model behavior in this harness is template-specific. The same harness, the same flag, two different models, two different responses. The flag is not a property of the API surface; it is a property of each template's Jinja-templated config. The harness has to assume the flag does nothing and provision budget.
Image-case perception gaps
Five generated-invoice cases. The first two passed the hidden oracle. p03 and p05 failed on perception-of-evidence rather than budget.
p03 includes previous_invoices.csv in the input folder. The expected audit flags INV-7801 for review with the warning code duplicate_risk because that invoice number already appeared in a previous batch. Bonsai approved INV-7801 instead. The cascade: total_approved_gross_cents shifted from 18 737 (ground truth) to 37 474 — 18737 × 2 = 37474, consistent with one duplicated invoice. proof_code propagated the difference. The Muse-Glimmer run made the same mistake on the same case prompt, on a slightly different invoice. Both models accepted the obvious-looking invoice without checking the prior-batch list.
p05 is a one-scan contact sheet case. The expected ignored_document_ids is ["QT-5601"]. Bonsai produced ["QUOTE QT-5601"] — the human-readable label from the document header, not the document ID printed in the bottom-right corner. The visible-checks rubric compares the value character-for-character; "QUOTE QT-5601" ≠ "QT-5601". The model parsed the wrong field as the document identifier. Across the other four paper cases where Bonsai's ignored_document_ids passed the rubric (QT-6400, CN-10032, CN-10032+QT-6400, ST-4170), it picked the right field. p05 is the case where the obvious-looking label is the wrong one.
Workflow shape errors
The four workflow cases (w04–w07) run a loop where the model picks actions like "read file" or "write artifact" and the harness executes them. Three of the four cases lost on shape compliance — the model wrote files in shapes the prompt did not specify.
w04 expects a normalized manifest wrapped in a normalized_manifest key:
{
"normalized_manifest": {
"case_id": "P3-WORK-04",
"active_files": [...],
"ignored_files": [...],
"normalized_files": [...]
}
}Bonsai wrote a flat object instead, with the four fields at the top level. The case prompt defines the schema; the harness's hidden oracle wraps one level deeper. The model read the prompt and wrote the file anyway. This is the same mistake Muse-Glimmer made on w04. The warnings_by_invoice["INV-9109"] also failed the hidden oracle: Bonsai produced ["missing_po", "inactive_vendor", "payment_short"], ground truth is ["inactive_vendor", "missing_payment", "missing_po"] — different codes (payment_short ≠ missing_payment) and different ordering.
w05 is the most error-dense case in the run. The expected workflow: approve INV-2204-R1, ignore the chat hint and proforma estimate, reject nothing. Bonsai approved INV-2204-R1 correctly but rejected INV-2204 (the original, superseded invoice) instead of ignoring it — INV-2204 ended up in reject_invoice_ids rather than ignored_document_ids. The evidence field missed three of six expected files. The proof_code was wrong: Bonsai wrote 47 268, ground truth is 47 825. Multi-error across invoice classification, evidence list, proof_code, and document selection.
w06 is the cleanest workflow case Bonsai produced. Core-oracle clean: both invoices (INV-3301, INV-3302) correctly identified, total split to 29 730 cents, proforma (PRO-3303) correctly ignored, proof code correct. The hidden oracle failed only on evidence:
incoming/attachments/invoice_3301.png
incoming/attachments/invoice_3302.png
incoming/bank_export_final.csv
incoming/attachments/remittance_advice.pngExpected:
incoming/attachments/invoice_3301.png
incoming/attachments/invoice_3302.png
incoming/attachments/remittance_advice.png
incoming/bank_export_final.csv
incoming/purchase_orders.csv
incoming/vendor_master.csvThe model included the four files it directly read and stopped. purchase_orders.csv and vendor_master.csv are listed in the prompt as available inputs but Bonsai treated "contributed to the audit" as "I literally opened this file" rather than "the prompt mentioned it as input".
w07 is a different failure mode from w04–w06: no output at all. The first run produced a loop where the model picked actions like "read file" or "write artifact" that ran for more than 30 minutes and was killed without producing audit_result.json or proof.txt. A retry with max_tokens: 4000 and the same template-kwargs flag ran for the full 1 800-second watchdog and ended with required_artifact_missing and no_output.
Where this puts the leaderboard
The nine-case paperwork-v3 suite uses two benchmarks: paperwork (five generated-invoice cases, scored as one aggregate run) and paperwork-workflow (four cases, scored individually). Bonsai now has both. Comparison rows:
| Rank | Model | practical_score | strict | core |
|---|---|---|---|---|
| 1 | opencode/minimax-m3-free | 88.9% | 8/9 | 8/9 |
| 4 | meta/muse-glimmer | 77.8% | 6/9 | 8/9 |
| 7 | google/gemma-4-26b-a4b | 61.1% | 4/9 | 7/9 |
| 8 | qwen3.8-27b | 55.6% | 3/9 | 7/9 |
| 9 | qwen3.6-35b-a3b | 38.9% | 1/9 | 6/9 |
| 10 | prism-ml/bonsai-27b | 33.3% | 2/9 | 4/9 |
| 11 | qwen3.6-flash | 33.3% | 0/9 | 6/9 |
Bonsai enters the leaderboard at 33.3% (slot 10 overall), winning the tie-breaker against Qwen3.6-Flash on strict resolved cases (2/9 vs 0/9). The gap to Muse-Glimmer (77.8%) is the difference between a model that holds the line on the workflow cases and one that loses them to shape errors: a flat manifest, a filename-as-document-ID, an invoice moved from ignored to rejected, two missing CSVs in an evidence list. The gap to gemma-4-26b-a4b (61.1%) shows higher strict (4/9) and core (7/9) rates where audit facts hold up even when artifact shapes wobble.
What the harness needs
The Muse-Glimmer article pointed at two harness changes: make max_tokens a per-model parameter, and add chat_template_kwargs as a first-class config option. The Bonsai run confirms both, and adds a third.
The third is prompt-comprehension robustness. Two reasoning models in a row have skipped the normalized_manifest wrapper key on w04 despite the case prompt defining the schema explicitly. Two reasoning models in a row have flagged the wrong field on p03's previous_invoices.csv check. One reasoning model read filenames as document IDs on p05. The bench writes prompts that assume the model reads the prompt carefully and writes the file exactly as specified. Reasoning models in this harness do not. The honest fix is to accept multiple equivalent artifact shapes — flat manifest OR wrapped manifest, label OR ID for ignored_document_ids — and score the substance rather than the shape.
The budget knobs are necessary. They are not sufficient. Bonsai 27B on paperwork-v3 is not a 2/9 model because of budget: the budget was right (8 000 tokens per image case, full watchdog on the workflow cases). Bonsai is a 2/9 model because the prompt asks for specific output shapes on specific cases, and a 27B reasoning model in this harness will skip the wrapper key, parse the wrong field, miss the credit-memo, and refuse to start the action loop on the credit-offset case. None of those are 2 400-token-empty-content failures. They are 8 000-token-clean-JSON-still-wrong failures.