Refining paperwork-v3 based on the Bonsai 27B run — Local Model Bench

What we found going back through the Bonsai 27B verify-output files case by case, and the bench-side changes that follow from each finding.

Local Model Bench editorial hero for Refining paperwork-v3. Title right: REFINING PAPERWORK-V3, three failure categories. Left: JSON snippets with sticky notes on the keys.

The Bonsai 27B paperwork run landed at 0 of 9 strict, 4 of 9 core_oracle, and 22.2% practical, putting it at rank 17 on the leaderboard. Strict-resolved means all four of the checks (audit_result_exists, visible_checks_pass, core_oracle_pass, hidden_oracle_pass) return true on a single case. Bonsai failed at least one check on every one of the nine cases. Core-oracle passed on four. The practical score is 0.5·(0/9) + 0.5·(4/9) = 0.222.

Looking at the leaderboard number tells you Bonsai did badly on paperwork-v3. Reading the per-case verify-output JSON tells you why, and gives the bench something to do about most of it. This post goes through what each case actually shows, and what changes to paperwork-v3 follow.

P01 and P02 — what the four core-passes got right

P01 and P02 are the only two paper cases that pass hidden_oracle_pass and core_oracle_pass. Both fail only on visible_checks_pass. The audit content is correct on both.

Looking at P01 against its ground_truth.json: Bonsai's audit_result.json matches on every field. case_id P3-GEN-01, approved_invoice_ids INV-7801, review_invoice_ids INV-7802 and INV-8422, reject_invoice_ids empty, ignored_document_ids QT-6400, total_approved_gross_cents 18,737, warnings_by_invoice, evidence list, proof_code 42,956 — all match. Yet the visible_checks rubric rejects the audit. The visible check is failing on something the audit content does not expose to a human reader.

P02 against ground_truth.json: case_id P3-GEN-02, approved_invoice_ids INV-82415, review_invoice_ids INV-82478, reject_invoice_ids INV-82533, ignored_document_ids CN-10032, total_approved_gross_cents 18,737, warnings_by_invoice, proof_code 266,454 — all match. The evidence list matches in content but the order differs: Bonsai wrote vendor_master.csv at index 2, ground_truth has it at index 6. That alone does not fail hidden_oracle_pass (the hidden oracle compares sorted contents), but it may explain the visible_checks failure.

P01 and P02 are the only two paper cases in the four core-oracle passes. The other two core-oracle passes are P04 and W06. P01 and P02 are also the only two cases where the audit content matches ground_truth exactly on every visible field — including the proof_code. The visible_checks rubric is the part of the bench that is failing these cases, and the rubric's failure is what makes strict=0/9 for the run.

P03 — the model skipped a step

The input folder for the third paper case includes a previous_invoices.csv. The ground-truth audit flags INV-7801 for review with the warning code duplicate_risk because that invoice number appeared in a previous batch. Bonsai approved INV-7801 instead, with no warnings on it. The expected audit has INV-7801 in review_invoice_ids; Bonsai put it in approved_invoice_ids. The cascade shows up in the totals: expected total_approved_gross_cents is 18,737. Bonsai emits 37,474, which is 18,737 × 2 exactly. The verify output tags this with the failure type duplicate_risk_missed as the primary cause, followed by total_calculation_error and proof_code_error.

This is a model perception failure. The CSV was in the input. The model did not read it the way the bench expects. We have Muse-Glimmer's P03 verify output on disk for comparison: Muse also approved INV-7801, with the same 37,474 total and the same duplicate_risk_missed failure type. Two reasoning models failing the same case prompt the same way is a prompt-clarity signal more than a capability ceiling. The paper cases will need prompts that make the previous-batch list more obviously part of the audit contract, not a side detail.

P05 — model picked the visible label over the printed ID

The expected ignored_document_ids for this case is ["QT-5601"]. Bonsai emitted ["QUOTE QT-5601"]. The visible-checks rubric does a character-for-character compare and rejects the label version. Failure type: ignored_document_id_error. On the other four paper cases, where Bonsai picked the right field, it passed this exact check. P05 is the case where the obvious-looking label is the wrong one, and the model went for the obvious one.

This is a smaller version of the same kind of failure as P03: the model is parsing the document but not parsing it the way the bench expects. A stricter prompt that names the printed ID field as the source would help. The bench-side fix would be to accept a substring match or a list alias, but that loses a real signal — these paper cases are designed to test which field the model reaches for.

P04 — substance right, evidence list flagged

Bonsai's P04 audit matches the expected on case_id, approved_invoice_ids, review_invoice_ids, reject_invoice_ids, and ignored_document_ids. core_oracle_pass returns true. hidden_oracle_pass is false on the single failure type missing_or_wrong_evidence. Looking at the actual evidence lists: Bonsai wrote six file paths (bank_export.csv, purchase_orders.csv, scans/INV-4170.png, scans/INV-4171.png, scans/ST-4170.png, vendor_master.csv). Expected wrote four paths: the same bank_export.csv, purchase_orders.csv, and vendor_master.csv, plus a single composite scans/orion_tax_collision_contact_sheet.png. Bonsai broke the composite contact-sheet into three underlying invoice scans (INV-4170.png, INV-4171.png, ST-4170.png); expected listed the contact sheet as one item. The substance of the audit is right; the path list structure is wrong. The rubric flagged this as missing_or_wrong_evidence without further detail; we have inferred the structural mismatch from comparing the lists ourselves.

W04 — structural fields match, warning codes do not

W04's case_id, approved_invoice_ids, review_invoice_ids, reject_invoice_ids, ignored_document_ids, total_approved_gross_cents, and proof_code all match the expected output exactly. The failure types are manifest_error and warning_code_error. We had to look at warnings_by_invoice to see what the warning_code mismatch actually is. For INV-9109 the expected warning list is ['inactive_vendor', 'missing_payment', 'missing_po']. Bonsai wrote ['missing_po', 'inactive_vendor', 'payment_short']. The codes are wrong: Bonsai used payment_short where expected was missing_payment. The order is also different. For INV-9108 the expected warning list is ['payment_short'] and Bonsai wrote ['payment_short'] — that one matches. The rubric flagged the warning code on INV-9109 plus an unspecified manifest check on the audit_result.json content.

W05 — seven failure types on one case

W05 is the case where the model and the rubric fail together. Seven failure types listed in the verify output: final_document_set_error, ignored_document_id_error, invoice_classification_error, missing_or_wrong_evidence, proof_code_error, proof_txt_error, wrong_document_selected.

Specifically: Bonsai approved INV-2204-R1 correctly but rejected INV-2204 (the original, superseded invoice) instead of ignoring it. INV-2204 ended up in reject_invoice_ids rather than ignored_document_ids. The expected audit has INV-2204 in ignored_document_ids. The proof code Bonsai emitted is 47,268; the expected is 47,825. The difference is 557 cents. The underlying cause is that the model classified the wrong invoice, which shifted the reconciled total, which shifted the proof code.

Bonsai's evidence list has four entries; expected has six. The two missing files are incoming/email_thread.txt and incoming/attachments/chat_hint.png. Bonsai omitted the email-thread file that drives the supersede-decision and the chat-hint screenshot that flags the credit memo.

A cleaner invoice-classification check that ignores the proof-code hash would still catch the model error. A looser evidence-list check would not — Bonsai's evidence list misses two of six expected files. So W05 is mixed: model errors on classification, rubric strictness on the path list.

W06 — only the evidence list

W06 is the cleanest failure in the run to diagnose. core_oracle_pass is true. hidden_oracle_pass is false, with the single failure type missing_or_wrong_evidence. Bonsai's proof code is 37,104 and matches the expected 37,104 exactly. Both invoices correctly identified. The total split correctly. The proforma correctly ignored. The only thing wrong is the evidence list — Bonsai wrote four file paths, expected six. The two missing are incoming/purchase_orders.csv and incoming/vendor_master.csv, the standard procurement reference files. Substantively the audit is right. Rubric-strictness on the path list is what failed.

W07 — no output

The verify output for W07 lists five failure types: final_document_set_error, no_output, payment_reconciliation_error, proof_txt_error, required_artifact_missing. The Bonsai run produced no audit_result.json and no proof.txt.

Looking at the agent_trace.json for this case: 24 trace steps. The agent attempted seven write_file actions, five read_file actions, two mkdir calls, and one list_files. Steps 1–15 produced parsed-and-acknowledged actions (read_file, list_files, mkdir, write_file). From step 16 onward, the harness returned parse failures on every emitted action. Steps 16 through 22 returned the same error: Invalid JSON action: Expecting property name enclosed in double quotes: line 1 column 275 (char 274). Steps 23 and 24 returned a different error: Invalid JSON action: Extra data: line 1 column 53 (char 52) and Invalid JSON action: Extra data: line 3 column 1 (char 314) — the model switched away from the property-name escaping approach and tried a different emission that also failed to parse. The trace ends at step 24 with no successful commit. No final artifact was written.

The watchdog for the Bonsai run is 1,800 seconds (the harness_metadata.watchdog_timeout_s field reads 1,800 in the scored JSON for this case). We cannot, on the data we have, distinguish a hang from a parse error from a budget exhaustion on a no-output result. The bench does not currently record the time between the harness dispatching the model's final commit and the harness receiving the result back. W07 is a failure with a known shape and an unknown cause.

What changes on the bench

The bench-side changes that fall out of the case failures above are specific to each category.

For W07, instrument the no-output path. The fix is bookkeeping on the harness side: timestamp the start of the model's final commit and the end. If the gap is short, that is a parse error. If the start never happened, that is a hang or pre-completion model stop. If the start happened but the end did not within the watch-dog, that is a budget exhaustion. Each cause has a different mitigation.

For W04, capture the rubric check that flagged the manifest against the actual audit_result.json content, so we know which field the rubric read that the human-readable field name did not surface.

For P03, change the paper-case prompt so the previous-batch list is named explicitly as part of the audit contract, not a side reference.

There is no single fix that would have moved Bonsai from 4 of 9 core to a higher score. P03 and P05 are model perception failures where the prompt is the right place to look first. P01 and P02 show that the model can write a clean audit — the rubric is failing them on the visible_checks shape comparison, not the model. W04, W05, and W06 are rubric-strictness failures where the rubric itself has to decide what it is measuring. W07 is a bench-instrumentation failure that has nothing to do with the model. The work for paperwork-v3 is to make each of those categories measurable separately, so the next low-scoring run can be attributed cleanly rather than read off a leaderboard.