Kimi K3 Max tops paperwork-v3 via Devin CLI: SWE-2 beats GPT-6 and Opus
Seven frontier model configs ran all nine paperwork-v3 cases through Devin CLI's agent loop. Kimi K3 Max takes first place on the board (94.4%), SWE-2 Max second (88.9%), ahead of GPT-6 Astra Max, Claude Opus 5.5 Max and Grok 4.7 XHigh.
We ran all nine paperwork-v3 cases through Devin CLI — five generated-scan audits plus the four workflow cases, same oracle — on seven frontier model configurations: GPT-6 Luna (high and max thinking), GPT-6 Astra Max, Claude Opus 5.5 Max, SWE-2 Max, Kimi K3 Max and Grok 4.7 XHigh. Kimi K3 Max finished first on the overall board; SWE-2 Max landed second, above every GPT-6 configuration and Claude Opus 5.5 Max.
Local Model Bench normally measures local models on Apple Silicon. This run is different: every model executed inside Devin CLI's agent loop — the model reads the scanned PNGs and CSVs with tools, writes audit_result.json, and a hidden oracle scores it exactly. That makes these rows comparable to our other agentic-harness rows (codex-default, opencode runs), not to the single-shot LM Studio rows.
The scoreboard
Each generated-scan case produces four checks: audit artifact written, visible invoice set correct, core oracle pass (right audit facts, cosmetic failures allowed), strict pass (exact match including evidence paths and proof code). Scan practical = 0.5 × (strict/5) + 0.5 × (core/5). The workflow cases score strict pass/fail plus a core-oracle check per case. The board column is the leaderboard's combined practical across all nine cases: 0.5 × (strict tasks/9) + 0.5 × (core tasks/9).
| Model (Devin CLI) | Scan strict | Scan core | Scan practical | Wf strict | Wf core | Board | Scan time/case |
|---|---|---|---|---|---|---|---|
| Kimi K3 Max | 4/5 | 5/5 | 90% | 4/4 | 4/4 | 94.4% (#1) | ~25 s |
| SWE-2 Max | 4/5 | 4/5 | 80% | 4/4† | 4/4 | 88.9% (#2) | ~64 s |
| GPT-6 Luna High | 3/5 | 3/5 | 60% | 4/4 | 4/4 | 77.8% (#5) | ~46 s |
| Claude Opus 5.5 Max | 3/5 | 3/5 | 60% | 4/4 | 4/4 | 77.8% (#6) | ~196 s |
| Grok 4.7 XHigh | 2/5 | 3/5 | 50% | 4/4 | 4/4 | 72.2% (#9) | ~84 s |
| GPT-6 Astra Max | 3/5 | 3/5 | 60% | 3/4 | 4/4 | 72.2% (#10) | ~97 s |
| GPT-6 Luna Max | 3/5 | 3/5 | 60% | 2/4 | 3/4 | 61.1% (#14) | ~61 s |
† SWE-2 ran workflow case P3-WORK-04 twice: one clean pass, one repeat run with a manifest slip; the board keeps the pass. Time column is the scan-suite average.
For reference, the existing best row on the generated-scan suite is GPT-5.4 Mini at 5/5 strict — run under a different harness — and the best local rows sit at 4/5 strict (Gemma-4 26B, Muse-Glimmer, MiniMax-M3, Bonsai 2 27B thinking).
The trap that caught six of seven
Case P3-GEN-02 contains a synthetic BrightPath invoice, INV-82533, stamped VENDOR HOLD / INACTIVE VENDOR with a red MISSING PO field. The bank export holds a matching row for it: 23,794 cents — exactly the invoice gross — with status pending. The task spec is unambiguous: payment_short applies when a paid bank row is lower than gross. A pending row is not a paid row, and 23,794 is not lower than 23,794.
Every GPT-6 configuration — Luna High, Luna Max, Astra Max — added payment_short to INV-82533 anyway, in both cases where the invoice appears. The wrong answers even share proof codes: all three returned 266551 on P3-GEN-02 (expected 266454), and two of three returned 290867 on P3-GEN-03 (expected 290770; Luna High returned 290964). That is what a systematic misreading looks like: all three treat "no paid bank row" as "paid zero". Claude Opus 5.5 Max made the identical error. So did SWE-2 Max — on P3-GEN-02 only; in the noisier eight-scan P3-GEN-03 it read the same pending row correctly.
Kimi's one arithmetic miss
Kimi K3 Max was the only configuration that never fell for the trap. Its single miss is arithmetic: on P3-GEN-01 it returned proof_code 43356 instead of 42956 — every classification, warning, ignored-document ID and evidence path exact. Under our scoring that still costs the strict check, so it lands at 4/5 strict and 5/5 core. It was between just under twice and about eight times faster per case than the rest of the field.
Grok's different failure
Grok 4.7 XHigh is the only model whose misses were not just the trap. On P3-GEN-02 it made the same pending-row misreading as the GPT-6 family. On P3-GEN-03 — the mixed eight-scan folder combining two earlier cases — it attached duplicate_risk to the wrong invoice (INV-82415 instead of INV-7801) and approved INV-7801 outright, a document-reading failure the other models did not make. On P3-GEN-01 it missed strict on a proof-code slip with all audit facts correct.
The workflow cases
The four workflow cases (messy intake, email-attachment intake, remittance split, credit offset) produced fewer failures than the scan suite: five of seven configurations swept all four. Grok 4.7 — weakest on the scan suite — went a clean 4/4 here — every one of its misses sits on the generated-scan suite, including the P3-GEN-03 duplicate-risk misattribution. The remaining misses: Astra dropped the credit-offset case (P3-WORK-07) on a final document-set error, SWE-2 ran the messy-intake case (P3-WORK-04) twice — one clean pass, one manifest slip on the repeat run — and Luna Max failed P3-WORK-04 and P3-WORK-05, the latter including the only core-oracle miss of the whole workflow batch, a wrong document selection. The max-thinking Luna tier scored below the High tier: more budget did not help.
The leaderboard
Seven new rows are on the overall board, marked as Devin CLI harness runs. Kimi K3 Max takes first place overall at 94.4% practical (8/9 strict) and SWE-2 Max ties opencode/minimax-m3-free at second (88.9%, 8/9), outscoring every GPT-6 configuration and Claude Opus 5.5 Max on this suite. The caveat is the sample: nine cases, one run each (one repeat for SWE-2's P3-WORK-04). The workflow results do as much of the ranking work as the scan trap — Luna Max's 2/4 there is what drops it to last of these seven rows — so read the ordering as a finding about pending-row reading plus workflow slips, not a verdict on overall capability.
Where it worked
- All seven configurations produced a valid
audit_result.jsonon every case — no extraction failures, no malformed JSON. - Invoice classification on the scan suite was correct on 34 of 35 runs; only Grok's P3-GEN-03 run misclassified.
- The Devin CLI harness read every scan image through its file tools; no vision-attachment workarounds were needed.
- Kimi K3 Max solved the pending-status trap on both cases it appears in, and five of seven configurations swept all four workflow cases.
- Grok 4.7 XHigh — weakest on scans — went 4/4 on the workflow cases; all its misses sit on the generated-scan suite.
Where it failed
- GPT-6 (all three tiers) and Claude Opus 5.5 Max added
payment_shortto the pending-payment invoice in P3-GEN-02 and P3-GEN-03 — a systematic family-level reading of "no paid row" as "paid zero". - SWE-2 Max made the same error on P3-GEN-02 but not on P3-GEN-03; on one run per case that is noise, not a property.
- Grok 4.7 XHigh misattributed
duplicate_riskin the mixed case, hit the pending trap on P3-GEN-02, and slipped a proof code on the simplest case. - GPT-6 Luna Max — the max-thinking tier — failed two of four workflow cases, including the batch's only core-oracle miss.
- None of the new runs reached GPT-5.4 Mini's existing 5/5 scan-suite reference row.
What was actually tested
- All nine paperwork-v3 cases: the five generated-scan cases (P3-GEN-01 to 05) plus the four workflow cases (P3-WORK-04 to 07).
- Harness: Devin CLI 3000.11.3 in non-interactive print mode, bypass permissions, disposable workspace copies containing only the visible case files. One run per case, except SWE-2's P3-WORK-04 which ran twice (one pass, one manifest slip).
- Scoring: hidden oracle exact match; scan cases score strict + core on nine output keys, workflow cases score strict on artifacts and audit plus a core-oracle check per case.
- These rows measure model-plus-harness; they are not comparable to raw-API rows on the same suite.
Verdict
On the nine-case paperwork benchmark through an agent harness, Kimi K3 Max takes first place overall (94.4%) and SWE-2 Max lands second (88.9%) — above Claude Opus 5.5 Max and every GPT-6 configuration, including Astra Max (72.2%) and Grok 4.7 XHigh (72.2%). Two things do the ranking work: a single scan-suite trap — a pending bank row equal to the invoice gross — that six of seven configurations misread at least once (systematically for the GPT-6 family and Opus), and workflow slips that split the 60%-scan cluster apart.