> ## Content Index
> Fetch the complete content index at: https://localmodelbench.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Refining paperwork-v3 based on the Bonsai 27B run — Local Model Bench
- URL: https://localmodelbench.com/refining-paperwork-v3-bonsai-27b/
- Published: 2026-09-21T16:23:12.000Z
- Updated: 2026-09-21T16:24:17.000Z
- Description: What we found going back through the Bonsai 27B verify-output files case by case, and the bench-side changes that follow from each finding.
- Author: LMB Editorial

The Bonsai 27B paperwork run landed at 0 of 9 strict, 4 of 9 core\_oracle, and 22.2% practical, putting it at rank 17 on the leaderboard. Strict-resolved means all four of the checks (`audit_result_exists`, `visible_checks_pass`, `core_oracle_pass`, `hidden_oracle_pass`) return true on a single case. Bonsai failed at least one check on every one of the nine cases. Core-oracle passed on four. The practical score is 0.5·(0/9) + 0.5·(4/9) = 0.222.

Looking at the leaderboard number tells you Bonsai did badly on paperwork-v3\. Reading the per-case verify-output JSON tells you why, and gives the bench something to do about most of it. This post goes through what each case actually shows, and what changes to paperwork-v3 follow.

## P01 and P02 — what the four core-passes got right

P01 and P02 are the only two paper cases that pass `hidden_oracle_pass` and `core_oracle_pass`. Both fail only on `visible_checks_pass`. The audit content is correct on both.

Looking at P01 against its ground\_truth.json: Bonsai's `audit_result.json` matches on every field. case\_id P3-GEN-01, approved\_invoice\_ids INV-7801, review\_invoice\_ids INV-7802 and INV-8422, reject\_invoice\_ids empty, ignored\_document\_ids QT-6400, total\_approved\_gross\_cents 18,737, warnings\_by\_invoice, evidence list, proof\_code 42,956 — all match. Yet the visible\_checks rubric rejects the audit. The visible check is failing on something the audit content does not expose to a human reader.

P02 against ground\_truth.json: case\_id P3-GEN-02, approved\_invoice\_ids INV-82415, review\_invoice\_ids INV-82478, reject\_invoice\_ids INV-82533, ignored\_document\_ids CN-10032, total\_approved\_gross\_cents 18,737, warnings\_by\_invoice, proof\_code 266,454 — all match. The evidence list matches in content but the order differs: Bonsai wrote `vendor_master.csv` at index 2, ground\_truth has it at index 6\. That alone does not fail `hidden_oracle_pass` (the hidden oracle compares sorted contents), but it may explain the visible\_checks failure.

P01 and P02 are the only two paper cases in the four core-oracle passes. The other two core-oracle passes are P04 and W06\. P01 and P02 are also the only two cases where the audit content matches ground\_truth exactly on every visible field — including the proof\_code. The visible\_checks rubric is the part of the bench that is failing these cases, and the rubric's failure is what makes strict=0/9 for the run.

## P03 — the model skipped a step

The input folder for the third paper case includes a `previous_invoices.csv`. The ground-truth audit flags `INV-7801` for review with the warning code `duplicate_risk` because that invoice number appeared in a previous batch. Bonsai approved `INV-7801` instead, with no warnings on it. The expected audit has `INV-7801` in `review_invoice_ids`; Bonsai put it in `approved_invoice_ids`. The cascade shows up in the totals: expected `total_approved_gross_cents` is 18,737\. Bonsai emits 37,474, which is 18,737 × 2 exactly. The verify output tags this with the failure type `duplicate_risk_missed` as the primary cause, followed by `total_calculation_error` and `proof_code_error`.

This is a model perception failure. The CSV was in the input. The model did not read it the way the bench expects. We have Muse-Glimmer's P03 verify output on disk for comparison: Muse also approved `INV-7801`, with the same 37,474 total and the same `duplicate_risk_missed` failure type. Two reasoning models failing the same case prompt the same way is a prompt-clarity signal more than a capability ceiling. The paper cases will need prompts that make the previous-batch list more obviously part of the audit contract, not a side detail.

## P05 — model picked the visible label over the printed ID

The expected `ignored_document_ids` for this case is `["QT-5601"]`. Bonsai emitted `["QUOTE QT-5601"]`. The visible-checks rubric does a character-for-character compare and rejects the label version. Failure type: `ignored_document_id_error`. On the other four paper cases, where Bonsai picked the right field, it passed this exact check. P05 is the case where the obvious-looking label is the wrong one, and the model went for the obvious one.

This is a smaller version of the same kind of failure as P03: the model is parsing the document but not parsing it the way the bench expects. A stricter prompt that names the printed ID field as the source would help. The bench-side fix would be to accept a substring match or a list alias, but that loses a real signal — these paper cases are designed to test which field the model reaches for.

## P04 — substance right, evidence list flagged

Bonsai's P04 audit matches the expected on case\_id, approved\_invoice\_ids, review\_invoice\_ids, reject\_invoice\_ids, and ignored\_document\_ids. `core_oracle_pass` returns true. `hidden_oracle_pass` is false on the single failure type `missing_or_wrong_evidence`. Looking at the actual evidence lists: Bonsai wrote six file paths (`bank_export.csv`, `purchase_orders.csv`, `scans/INV-4170.png`, `scans/INV-4171.png`, `scans/ST-4170.png`, `vendor_master.csv`). Expected wrote four paths: the same `bank_export.csv`, `purchase_orders.csv`, and `vendor_master.csv`, plus a single composite `scans/orion_tax_collision_contact_sheet.png`. Bonsai broke the composite contact-sheet into three underlying invoice scans (`INV-4170.png`, `INV-4171.png`, `ST-4170.png`); expected listed the contact sheet as one item. The substance of the audit is right; the path list structure is wrong. The rubric flagged this as `missing_or_wrong_evidence` without further detail; we have inferred the structural mismatch from comparing the lists ourselves.

## W04 — structural fields match, warning codes do not

W04's case\_id, approved\_invoice\_ids, review\_invoice\_ids, reject\_invoice\_ids, ignored\_document\_ids, total\_approved\_gross\_cents, and proof\_code all match the expected output exactly. The failure types are `manifest_error` and `warning_code_error`. We had to look at warnings\_by\_invoice to see what the warning\_code mismatch actually is. For `INV-9109` the expected warning list is `['inactive_vendor', 'missing_payment', 'missing_po']`. Bonsai wrote `['missing_po', 'inactive_vendor', 'payment_short']`. The codes are wrong: Bonsai used `payment_short` where expected was `missing_payment`. The order is also different. For `INV-9108` the expected warning list is `['payment_short']` and Bonsai wrote `['payment_short']` — that one matches. The rubric flagged the warning code on INV-9109 plus an unspecified manifest check on the audit\_result.json content.

## W05 — seven failure types on one case

W05 is the case where the model and the rubric fail together. Seven failure types listed in the verify output: `final_document_set_error`, `ignored_document_id_error`, `invoice_classification_error`, `missing_or_wrong_evidence`, `proof_code_error`, `proof_txt_error`, `wrong_document_selected`.

Specifically: Bonsai approved `INV-2204-R1` correctly but rejected `INV-2204` (the original, superseded invoice) instead of ignoring it. `INV-2204` ended up in `reject_invoice_ids` rather than `ignored_document_ids`. The expected audit has `INV-2204` in `ignored_document_ids`. The proof code Bonsai emitted is 47,268; the expected is 47,825\. The difference is 557 cents. The underlying cause is that the model classified the wrong invoice, which shifted the reconciled total, which shifted the proof code.

Bonsai's evidence list has four entries; expected has six. The two missing files are `incoming/email_thread.txt` and `incoming/attachments/chat_hint.png`. Bonsai omitted the email-thread file that drives the supersede-decision and the chat-hint screenshot that flags the credit memo.

A cleaner invoice-classification check that ignores the proof-code hash would still catch the model error. A looser evidence-list check would not — Bonsai's evidence list misses two of six expected files. So W05 is mixed: model errors on classification, rubric strictness on the path list.

## W06 — only the evidence list

W06 is the cleanest failure in the run to diagnose. `core_oracle_pass` is true. `hidden_oracle_pass` is false, with the single failure type `missing_or_wrong_evidence`. Bonsai's proof code is 37,104 and matches the expected 37,104 exactly. Both invoices correctly identified. The total split correctly. The proforma correctly ignored. The only thing wrong is the evidence list — Bonsai wrote four file paths, expected six. The two missing are `incoming/purchase_orders.csv` and `incoming/vendor_master.csv`, the standard procurement reference files. Substantively the audit is right. Rubric-strictness on the path list is what failed.

## W07 — no output

The verify output for W07 lists five failure types: `final_document_set_error`, `no_output`, `payment_reconciliation_error`, `proof_txt_error`, `required_artifact_missing`. The Bonsai run produced no `audit_result.json` and no `proof.txt`.

Looking at the agent\_trace.json for this case: 24 trace steps. The agent attempted seven `write_file` actions, five `read_file` actions, two `mkdir` calls, and one `list_files`. Steps 1–15 produced parsed-and-acknowledged actions (read\_file, list\_files, mkdir, write\_file). From step 16 onward, the harness returned parse failures on every emitted action. Steps 16 through 22 returned the same error: `Invalid JSON action: Expecting property name enclosed in double quotes: line 1 column 275 (char 274)`. Steps 23 and 24 returned a different error: `Invalid JSON action: Extra data: line 1 column 53 (char 52)` and `Invalid JSON action: Extra data: line 3 column 1 (char 314)` — the model switched away from the property-name escaping approach and tried a different emission that also failed to parse. The trace ends at step 24 with no successful commit. No final artifact was written.

The watchdog for the Bonsai run is 1,800 seconds (the `harness_metadata.watchdog_timeout_s` field reads 1,800 in the scored JSON for this case). We cannot, on the data we have, distinguish a hang from a parse error from a budget exhaustion on a no-output result. The bench does not currently record the time between the harness dispatching the model's final commit and the harness receiving the result back. W07 is a failure with a known shape and an unknown cause.

## What changes on the bench

The bench-side changes that fall out of the case failures above are specific to each category.

For W07, instrument the no-output path. The fix is bookkeeping on the harness side: timestamp the start of the model's final commit and the end. If the gap is short, that is a parse error. If the start never happened, that is a hang or pre-completion model stop. If the start happened but the end did not within the watch-dog, that is a budget exhaustion. Each cause has a different mitigation.

For W04, capture the rubric check that flagged the manifest against the actual `audit_result.json` content, so we know which field the rubric read that the human-readable field name did not surface.

For P03, change the paper-case prompt so the previous-batch list is named explicitly as part of the audit contract, not a side reference.

There is no single fix that would have moved Bonsai from 4 of 9 core to a higher score. P03 and P05 are model perception failures where the prompt is the right place to look first. P01 and P02 show that the model can write a clean audit — the rubric is failing them on the visible\_checks shape comparison, not the model. W04, W05, and W06 are rubric-strictness failures where the rubric itself has to decide what it is measuring. W07 is a bench-instrumentation failure that has nothing to do with the model. The work for paperwork-v3 is to make each of those categories measurable separately, so the next low-scoring run can be attributed cleanly rather than read off a leaderboard.