> ## Content Index
> Fetch the complete content index at: https://localmodelbench.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Kimi K3 Max tops paperwork-v3 via Devin CLI: SWE-2 beats GPT-6 and Opus
- URL: https://localmodelbench.com/frontier-models-on-paperwork-v3-via-devin-cli/
- Published: 2026-09-27T16:13:05.000Z
- Updated: 2026-09-28T10:50:22.000Z
- Description: Seven frontier model configs ran all nine paperwork-v3 cases through Devin CLI's agent loop. Kimi K3 Max takes first place on the board (94.4%), SWE-2 Max second (88.9%), ahead of GPT-6 Astra Max, Claude Opus 5.5 Max and Grok 4.7 XHigh.
- Author: LMB Editorial
- Tags: paperwork-v3, devin-cli, comparison

We ran all nine paperwork-v3 cases through Devin CLI — five generated-scan audits plus the four workflow cases, same oracle — on seven frontier model configurations: GPT-6 Luna (high and max thinking), GPT-6 Astra Max, Claude Opus 5.5 Max, SWE-2 Max, Kimi K3 Max and Grok 4.7 XHigh. Kimi K3 Max finished first on the overall board; SWE-2 Max landed second, above every GPT-6 configuration and Claude Opus 5.5 Max.

Local Model Bench normally measures local models on Apple Silicon. This run is different: every model executed inside Devin CLI's agent loop — the model reads the scanned PNGs and CSVs with tools, writes `audit_result.json`, and a hidden oracle scores it exactly. That makes these rows comparable to our other agentic-harness rows (codex-default, opencode runs), not to the single-shot LM Studio rows.

## The scoreboard

Each generated-scan case produces four checks: audit artifact written, visible invoice set correct, core oracle pass (right audit facts, cosmetic failures allowed), strict pass (exact match including evidence paths and proof code). Scan practical = 0.5 × (strict/5) + 0.5 × (core/5). The workflow cases score strict pass/fail plus a core-oracle check per case. The board column is the leaderboard's combined practical across all nine cases: 0.5 × (strict tasks/9) + 0.5 × (core tasks/9).

| Model (Devin CLI)   | Scan strict | Scan core | Scan practical | Wf strict | Wf core | Board       | Scan time/case |
| ------------------- | ----------- | --------- | -------------- | --------- | ------- | ----------- | -------------- |
| Kimi K3 Max         | 4/5         | 5/5       | 90%            | 4/4       | 4/4     | 94.4% (#1)  | \~25 s         |
| SWE-2 Max           | 4/5         | 4/5       | 80%            | 4/4†      | 4/4     | 88.9% (#2)  | \~64 s         |
| GPT-6 Luna High     | 3/5         | 3/5       | 60%            | 4/4       | 4/4     | 77.8% (#5)  | \~46 s         |
| Claude Opus 5.5 Max | 3/5         | 3/5       | 60%            | 4/4       | 4/4     | 77.8% (#6)  | \~196 s        |
| Grok 4.7 XHigh      | 2/5         | 3/5       | 50%            | 4/4       | 4/4     | 72.2% (#9)  | \~84 s         |
| GPT-6 Astra Max     | 3/5         | 3/5       | 60%            | 3/4       | 4/4     | 72.2% (#10) | \~97 s         |
| GPT-6 Luna Max      | 3/5         | 3/5       | 60%            | 2/4       | 3/4     | 61.1% (#14) | \~61 s         |

† SWE-2 ran workflow case P3-WORK-04 twice: one clean pass, one repeat run with a manifest slip; the board keeps the pass. Time column is the scan-suite average.

For reference, the existing best row on the generated-scan suite is GPT-5.4 Mini at 5/5 strict — run under a different harness — and the best local rows sit at 4/5 strict (Gemma-4 26B, Muse-Glimmer, MiniMax-M3, Bonsai 2 27B thinking).

## The trap that caught six of seven

Case P3-GEN-02 contains a synthetic BrightPath invoice, INV-82533, stamped `VENDOR HOLD / INACTIVE VENDOR` with a red `MISSING PO` field. The bank export holds a matching row for it: 23,794 cents — exactly the invoice gross — with status `pending`. The task spec is unambiguous: `payment_short` applies when a *paid* bank row is lower than gross. A pending row is not a paid row, and 23,794 is not lower than 23,794.

Every GPT-6 configuration — Luna High, Luna Max, Astra Max — added `payment_short` to INV-82533 anyway, in both cases where the invoice appears. The wrong answers even share proof codes: all three returned 266551 on P3-GEN-02 (expected 266454), and two of three returned 290867 on P3-GEN-03 (expected 290770; Luna High returned 290964). That is what a systematic misreading looks like: all three treat "no paid bank row" as "paid zero". Claude Opus 5.5 Max made the identical error. So did SWE-2 Max — on P3-GEN-02 only; in the noisier eight-scan P3-GEN-03 it read the same pending row correctly.

## Kimi's one arithmetic miss

Kimi K3 Max was the only configuration that never fell for the trap. Its single miss is arithmetic: on P3-GEN-01 it returned proof\_code 43356 instead of 42956 — every classification, warning, ignored-document ID and evidence path exact. Under our scoring that still costs the strict check, so it lands at 4/5 strict and 5/5 core. It was between just under twice and about eight times faster per case than the rest of the field.

## Grok's different failure

Grok 4.7 XHigh is the only model whose misses were not just the trap. On P3-GEN-02 it made the same pending-row misreading as the GPT-6 family. On P3-GEN-03 — the mixed eight-scan folder combining two earlier cases — it attached `duplicate_risk` to the wrong invoice (INV-82415 instead of INV-7801) and approved INV-7801 outright, a document-reading failure the other models did not make. On P3-GEN-01 it missed strict on a proof-code slip with all audit facts correct.

## The workflow cases

The four workflow cases (messy intake, email-attachment intake, remittance split, credit offset) produced fewer failures than the scan suite: five of seven configurations swept all four. Grok 4.7 — weakest on the scan suite — went a clean 4/4 here — every one of its misses sits on the generated-scan suite, including the P3-GEN-03 duplicate-risk misattribution. The remaining misses: Astra dropped the credit-offset case (P3-WORK-07) on a final document-set error, SWE-2 ran the messy-intake case (P3-WORK-04) twice — one clean pass, one manifest slip on the repeat run — and Luna Max failed P3-WORK-04 and P3-WORK-05, the latter including the only core-oracle miss of the whole workflow batch, a wrong document selection. The max-thinking Luna tier scored below the High tier: more budget did not help.

## The leaderboard

Seven new rows are on the overall board, marked as Devin CLI harness runs. Kimi K3 Max takes first place overall at 94.4% practical (8/9 strict) and SWE-2 Max ties opencode/minimax-m3-free at second (88.9%, 8/9), outscoring every GPT-6 configuration and Claude Opus 5.5 Max on this suite. The caveat is the sample: nine cases, one run each (one repeat for SWE-2's P3-WORK-04). The workflow results do as much of the ranking work as the scan trap — Luna Max's 2/4 there is what drops it to last of these seven rows — so read the ordering as a finding about pending-row reading plus workflow slips, not a verdict on overall capability.

## Where it worked

- All seven configurations produced a valid `audit_result.json` on every case — no extraction failures, no malformed JSON.
- Invoice classification on the scan suite was correct on 34 of 35 runs; only Grok's P3-GEN-03 run misclassified.
- The Devin CLI harness read every scan image through its file tools; no vision-attachment workarounds were needed.
- Kimi K3 Max solved the pending-status trap on both cases it appears in, and five of seven configurations swept all four workflow cases.
- Grok 4.7 XHigh — weakest on scans — went 4/4 on the workflow cases; all its misses sit on the generated-scan suite.

## Where it failed

- GPT-6 (all three tiers) and Claude Opus 5.5 Max added `payment_short` to the pending-payment invoice in P3-GEN-02 and P3-GEN-03 — a systematic family-level reading of "no paid row" as "paid zero".
- SWE-2 Max made the same error on P3-GEN-02 but not on P3-GEN-03; on one run per case that is noise, not a property.
- Grok 4.7 XHigh misattributed `duplicate_risk` in the mixed case, hit the pending trap on P3-GEN-02, and slipped a proof code on the simplest case.
- GPT-6 Luna Max — the max-thinking tier — failed two of four workflow cases, including the batch's only core-oracle miss.
- None of the new runs reached GPT-5.4 Mini's existing 5/5 scan-suite reference row.

## What was actually tested

- All nine paperwork-v3 cases: the five generated-scan cases (P3-GEN-01 to 05) plus the four workflow cases (P3-WORK-04 to 07).
- Harness: Devin CLI 3000.11.3 in non-interactive print mode, bypass permissions, disposable workspace copies containing only the visible case files. One run per case, except SWE-2's P3-WORK-04 which ran twice (one pass, one manifest slip).
- Scoring: hidden oracle exact match; scan cases score strict + core on nine output keys, workflow cases score strict on artifacts and audit plus a core-oracle check per case.
- These rows measure model-plus-harness; they are not comparable to raw-API rows on the same suite.

## Verdict

> On the nine-case paperwork benchmark through an agent harness, Kimi K3 Max takes first place overall (94.4%) and SWE-2 Max lands second (88.9%) — above Claude Opus 5.5 Max and every GPT-6 configuration, including Astra Max (72.2%) and Grok 4.7 XHigh (72.2%). Two things do the ranking work: a single scan-suite trap — a pending bank row equal to the invoice gross — that six of seven configurations misread at least once (systematically for the GPT-6 family and Opus), and workflow slips that split the 60%-scan cluster apart.

## Sources

- [Local Model Bench leaderboard](https://localmodelbench.com/leaderboard/)
- [Paperwork Trial v3 suite](https://localmodelbench.com/paperwork-trial-v3/)
- [The no-think leaderboard that wasn't](https://localmodelbench.com/the-no-think-leaderboard-that-wasnt/)