> ## Content Index
> Fetch the complete content index at: https://localmodelbench.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Bonsai 27B on paperwork-v3: 2/9 resolved, 4/9 core (33.3% practical)
- URL: https://localmodelbench.com/bonsai-27b-paperwork-v3/
- Published: 2026-09-19T18:40:14.000Z
- Updated: 2026-09-23T18:29:02.000Z
- Description: prism-ml/bonsai-27b scored 2/9 strict and 4/9 core on paperwork-v3 (33.3% practical), ranking #10 on the leaderboard. Like Muse-Glimmer, it required an 8,000 max_tokens ceiling, but its llama.cpp template ignored reasoning flags. Failures were shape and evidence errors, not token stalls.
- Author: LMB Editorial

prism-ml/bonsai-27b is a 27B reasoning model loaded into LM Studio and called through the OpenAI-compatible endpoint at `http://localhost:1234/v1` with non-default sampling settings. On the nine-case paperwork-v3 suite it scored 2/9 strict and 4/9 core, for a practical\_score of 33.3%, ranking tenth on the public leaderboard. Same harness, same `enable_thinking: false` flag, same shape of initial failure as the Muse-Glimmer run — except the llama.cpp template bundled with Bonsai reads the flag and ignores it.

## What ran

prism-ml/bonsai-27b was loaded in LM Studio and called through the OpenAI-compatible endpoint at `http://localhost:1234/v1` with non-default sampling settings (`temperature: 0`, `top_p: 1`). Five generated-invoice cases (p01–p05) and four paperwork-workflow cases (w04–w07) ran sequentially against the paperwork-v3 bench, mirroring the harness configuration used for the Muse-Glimmer run.

| Case                                          | Type     | core\_ok           | hidden\_oracle       | elapsed | completion tokens | of which reasoning |
| --------------------------------------------- | -------- | ------------------ | -------------------- | ------- | ----------------- | ------------------ |
| p01 (generated\_invoice\_case\_01)            | image    | yes                | yes                  | \~6:00  | 5 563             | 5 307              |
| p02 (generated\_invoice\_case\_02)            | image    | yes                | yes                  | \~5:30  | 5 234             | 5 005              |
| **p03 (generated\_invoice\_case\_03)**        | image    | **no**             | **no**               | \~9:00  | 7 801             | **7 386**          |
| p04 (generated\_invoice\_case\_04)            | image    | yes                | **no (path format)** | \~6:30  | 5 889             | 5 612              |
| **p05 (generated\_invoice\_case\_05)**        | image    | **no**             | **no**               | \~5:30  | 5 122             | **4 937**          |
| **w04 (messy\_intake\_workflow\_case\_04)**   | workflow | **no**             | **no**               | \~8:00  | n/a               | n/a                |
| **w05 (email\_attachment\_intake\_case\_05)** | workflow | **no**             | **no**               | \~10:00 | n/a               | n/a                |
| w06 (remittance\_split\_case\_06)             | workflow | yes                | **no (evidence)**    | \~9:00  | n/a               | n/a                |
| **w07 (credit\_offset\_case\_07)**            | workflow | **no (no output)** | **no**               | timeout | n/a               | n/a                |

Two clean hidden-oracle passes (p01, p02), two core-only passes (p04, w06) where audit fields matched ground truth but an artifact shape or evidence list failed its hidden-oracle structural check, and five core fails. That counts as 2/9 strict (22.2%) and 4/9 core (44.4%), for an aggregate practical\_score of 0.5 × (2/9) + 0.5 × (4/9) = 6/18 = 33.3%.

## Bonsai's template ignored the reasoning-suppression flag

The Muse-Glimmer run established that LM Studio reasoning models in this harness need either a larger `max_tokens` ceiling or a working `chat_template_kwargs: {"enable_thinking": false}` flag. Muse-Glimmer's llama.cpp template respected the flag and stopped spending tokens on the hidden reasoning trace. Bonsai's template reads the request field and does nothing with it.

The diagnostic: a single Bonsai run with `max_tokens: 2400` and `chat_template_kwargs: {"enable_thinking": false}` produced the same shape of failure the Muse run did on the first attempt — `finish_reason: length`, `content: ""`, the reasoning trace eating the budget. Bumping `max_tokens` to 8 000 produced a clean JSON body. Sending the template-kwargs flag on top did not change the trace length.

| max\_tokens | template\_kwargs            | reasoning\_tokens | completion\_tokens | content   | finish\_reason | core\_ok |
| ----------- | --------------------------- | ----------------- | ------------------ | --------- | -------------- | -------- |
| 2 400       | {"enable\_thinking": false} | \~2 400           | 2 400              | ""        | length         | no       |
| 8 000       | {"enable\_thinking": false} | 5 307             | 5 563              | full JSON | stop           | yes      |

Several reasoning-suppression flags — the top-level `chat_template_kwargs` variant and the OpenAI-API equivalents — were all tested on Bonsai. None of them reduced the trace length. The only knob that did was `max_tokens`. Same shape as Muse-Glimmer's first run, but no alternative intervention to recover.

Reasoning-model behavior in this harness is template-specific. The same harness, the same flag, two different models, two different responses. The flag is not a property of the API surface; it is a property of each template's Jinja-templated config. The harness has to assume the flag does nothing and provision budget.

## Image-case perception gaps

Five generated-invoice cases. The first two passed the hidden oracle. p03 and p05 failed on perception-of-evidence rather than budget.

**p03** includes `previous_invoices.csv` in the input folder. The expected audit flags `INV-7801` for review with the warning code `duplicate_risk` because that invoice number already appeared in a previous batch. Bonsai approved `INV-7801` instead. The cascade: `total_approved_gross_cents` shifted from 18 737 (ground truth) to 37 474 — `18737 × 2 = 37474`, consistent with one duplicated invoice. `proof_code` propagated the difference. The Muse-Glimmer run made the same mistake on the same case prompt, on a slightly different invoice. Both models accepted the obvious-looking invoice without checking the prior-batch list.

**p05** is a one-scan contact sheet case. The expected `ignored_document_ids` is `["QT-5601"]`. Bonsai produced `["QUOTE QT-5601"]` — the human-readable label from the document header, not the document ID printed in the bottom-right corner. The visible-checks rubric compares the value character-for-character; "QUOTE QT-5601" ≠ "QT-5601". The model parsed the wrong field as the document identifier. Across the other four paper cases where Bonsai's `ignored_document_ids` passed the rubric (QT-6400, CN-10032, CN-10032+QT-6400, ST-4170), it picked the right field. p05 is the case where the obvious-looking label is the wrong one.

## Workflow shape errors

The four workflow cases (w04–w07) run a loop where the model picks actions like "read file" or "write artifact" and the harness executes them. Three of the four cases lost on shape compliance — the model wrote files in shapes the prompt did not specify.

**w04** expects a normalized manifest wrapped in a `normalized_manifest` key:

```
{
  "normalized_manifest": {
    "case_id": "P3-WORK-04",
    "active_files": [...],
    "ignored_files": [...],
    "normalized_files": [...]
  }
}
```

Bonsai wrote a flat object instead, with the four fields at the top level. The case prompt defines the schema; the harness's hidden oracle wraps one level deeper. The model read the prompt and wrote the file anyway. This is the same mistake Muse-Glimmer made on w04\. The `warnings_by_invoice["INV-9109"]` also failed the hidden oracle: Bonsai produced `["missing_po", "inactive_vendor", "payment_short"]`, ground truth is `["inactive_vendor", "missing_payment", "missing_po"]` — different codes (`payment_short` ≠ `missing_payment`) and different ordering.

**w05** is the most error-dense case in the run. The expected workflow: approve `INV-2204-R1`, ignore the chat hint and proforma estimate, reject nothing. Bonsai approved `INV-2204-R1` correctly but rejected `INV-2204` (the original, superseded invoice) instead of ignoring it — `INV-2204` ended up in `reject_invoice_ids` rather than `ignored_document_ids`. The `evidence` field missed three of six expected files. The `proof_code` was wrong: Bonsai wrote 47 268, ground truth is 47 825\. Multi-error across invoice classification, evidence list, proof\_code, and document selection.

**w06** is the cleanest workflow case Bonsai produced. Core-oracle clean: both invoices (`INV-3301`, `INV-3302`) correctly identified, total split to 29 730 cents, proforma (`PRO-3303`) correctly ignored, proof code correct. The hidden oracle failed only on `evidence`:

```
incoming/attachments/invoice_3301.png
incoming/attachments/invoice_3302.png
incoming/bank_export_final.csv
incoming/attachments/remittance_advice.png
```

Expected:

```
incoming/attachments/invoice_3301.png
incoming/attachments/invoice_3302.png
incoming/attachments/remittance_advice.png
incoming/bank_export_final.csv
incoming/purchase_orders.csv
incoming/vendor_master.csv
```

The model included the four files it directly read and stopped. `purchase_orders.csv` and `vendor_master.csv` are listed in the prompt as available inputs but Bonsai treated "contributed to the audit" as "I literally opened this file" rather than "the prompt mentioned it as input".

**w07** is a different failure mode from w04–w06: no output at all. The first run produced a loop where the model picked actions like "read file" or "write artifact" that ran for more than 30 minutes and was killed without producing `audit_result.json` or `proof.txt`. A retry with `max_tokens: 4000` and the same template-kwargs flag ran for the full 1 800-second watchdog and ended with `required_artifact_missing` and `no_output`.

## Where this puts the leaderboard

The nine-case paperwork-v3 suite uses two benchmarks: paperwork (five generated-invoice cases, scored as one aggregate run) and paperwork-workflow (four cases, scored individually). Bonsai now has both. Comparison rows:

| Rank   | Model                    | practical\_score | strict  | core    |
| ------ | ------------------------ | ---------------- | ------- | ------- |
| 1      | opencode/minimax-m3-free | 88.9%            | 8/9     | 8/9     |
| 4      | meta/muse-glimmer        | 77.8%            | 6/9     | 8/9     |
| 7      | google/gemma-4-26b-a4b   | 61.1%            | 4/9     | 7/9     |
| 8      | qwen3.8-27b              | 55.6%            | 3/9     | 7/9     |
| 9      | qwen3.6-35b-a3b          | 38.9%            | 1/9     | 6/9     |
| **10** | **prism-ml/bonsai-27b**  | **33.3%**        | **2/9** | **4/9** |
| 11     | qwen3.6-flash            | 33.3%            | 0/9     | 6/9     |

Bonsai enters the leaderboard at 33.3% (slot 10 overall), winning the tie-breaker against Qwen3.6-Flash on strict resolved cases (2/9 vs 0/9). The gap to Muse-Glimmer (77.8%) is the difference between a model that holds the line on the workflow cases and one that loses them to shape errors: a flat manifest, a filename-as-document-ID, an invoice moved from `ignored` to `rejected`, two missing CSVs in an evidence list. The gap to gemma-4-26b-a4b (61.1%) shows higher strict (4/9) and core (7/9) rates where audit facts hold up even when artifact shapes wobble.

## What the harness needs

The Muse-Glimmer article pointed at two harness changes: make `max_tokens` a per-model parameter, and add `chat_template_kwargs` as a first-class config option. The Bonsai run confirms both, and adds a third.

The third is prompt-comprehension robustness. Two reasoning models in a row have skipped the `normalized_manifest` wrapper key on w04 despite the case prompt defining the schema explicitly. Two reasoning models in a row have flagged the wrong field on p03's `previous_invoices.csv` check. One reasoning model read filenames as document IDs on p05\. The bench writes prompts that assume the model reads the prompt carefully and writes the file exactly as specified. Reasoning models in this harness do not. The honest fix is to accept multiple equivalent artifact shapes — flat manifest OR wrapped manifest, label OR ID for `ignored_document_ids` — and score the substance rather than the shape.

The budget knobs are necessary. They are not sufficient. Bonsai 27B on paperwork-v3 is not a 2/9 model because of budget: the budget was right (8 000 tokens per image case, full watchdog on the workflow cases). Bonsai is a 2/9 model because the prompt asks for specific output shapes on specific cases, and a 27B reasoning model in this harness will skip the wrapper key, parse the wrong field, miss the credit-memo, and refuse to start the action loop on the credit-offset case. None of those are 2 400-token-empty-content failures. They are 8 000-token-clean-JSON-still-wrong failures.