Qwen3.8 27B on paperwork-v3: 3/9 closed, down from Qwen3.6 27B's 5/9 on the same nine cases
Qwen3.8 27B is the successor to Qwen3.6 27B, which currently leads the Local Model Bench local row on the nine-case paperwork-v3 suite. On this generation Qwen3.8 closed 3/9 strictly and 7/9 core-oracle for a practical score of 55.6%. Qwen3.6 27B closed 5/9 strictly and 8/9 core-oracle for 72.2% on

Qwen3.8 27B is the successor to Qwen3.6 27B, which currently leads the Local Model Bench local row on the nine-case paperwork-v3 suite. On this generation Qwen3.8 closed 3/9 strictly and 7/9 core-oracle for a practical score of 55.6%. Qwen3.6 27B closed 5/9 strictly and 8/9 core-oracle for 72.2% on the same nine cases earlier in the year. With a single run per case, a one-task delta cannot be separated from run-to-run variation.
Qwen3.8 27B is the natural next candidate to test against Qwen3.6 27B on Local Model Bench. Both are dense open-weight Qwen models in the 27B class, both run locally on a Mac mini M4 with 64 GB unified memory, and both target the same closed-document paperwork workload. Qwen3.8 is the newer release. The question is whether it closes more of the boring last mile.
On this nine-case run, with one run per case, Qwen3.8 27B closed 3/9 strictly and 7/9 core-oracle for a practical score of 55.6%; Qwen3.6 27B closed 5/9 strictly and 8/9 core-oracle for 72.2% on the same nine cases earlier in the year. Both deltas are one task. The interesting part is where it closed, where it nearly closed, and where the pattern of failure suggests a different test shape rather than a strictly weaker model.
What 3/9 looks like at the case level
Local Model Bench separates two signals per case: a strict resolved signal (the model produced a final artifact the hidden oracle accepted as exact) and a core-oracle pass signal (the model identified the right audit facts and only failed on cosmetic or arithmetic closure). The combination of those two signals is the practical score.
On the five generated-invoice paper-trail cases, Qwen3.8 27B reached a core-oracle pass on three: P01 two-invoice PO revision, P02 two invoices with a credit note, and P05 three-vendor mix. None of those five cases closed strictly. Every single one failed on the proof-code arithmetic at the end of the case.
- P01: core-pass, failed on proof-code. Two-invoice folder with one PO revision. Core audit facts right, sum-arithmetic off.
- P02: core-pass, failed on proof-code. Credit note plus partial payment plus vendor hold. Same pattern.
- P03: failed on five checks. Three-invoice folder with a duplicate risk, a stale export, and a short payment. Model misclassified invoices and missed the duplicate warning.
- P04: failed on four checks. Mix of paid, disputed, and revised invoices. Model misclassified and missed totals.
- P05: core-pass, failed on proof-code. Three-vendor batch. Audit facts right, sum-arithmetic off.
The workflow cases: one clean, three messy
The four agentic workflow cases are different. They are not a single audit prompt. They are messy folders the model has to walk through: read the README, decide which files are active versus draft, preserve protected source folders, write intermediate artifacts and the final audit result, and produce a proof file with a numeric checksum.
Qwen3.8 27B closed one of those four workflows strictly. W05, the email-attachment intake case, was a clean run. The model selected the right email thread, identified the three invoices and the memo, classified them correctly, and produced a passing proof. That is the only fully resolved workflow case for Qwen3.8 27B in this run.
The other three workflow cases ranged from near misses to broad failures. W06 remittance-split reached core-oracle pass but did not produce the required final-document-set and normalized-text artifacts. W07 credit-offset failed on payment reconciliation and several artifact checks. W04 messy-intake broke almost everywhere: document index, ignored document IDs, invoice classification, manifest, normalized text, proof code, proof file, totals, and warning codes.
Why proof-code fails everywhere
The single most common failure for Qwen3.8 27B on this suite was proof_code_error. Every single one of the five generated-invoice cases failed on the same arithmetic step at the end. Three of the four workflow cases failed the same step too.
The paper-trail proof-code formula is deterministic: sum of approved invoice gross cents plus the sum of the numeric parts of every invoice ID in the approved, review, and reject lists plus ninety-seven times the total warning count. The model produced the right invoice classifications and the right totals often. The arithmetic wrapper at the end is the part that missed.
That is exactly the boring last mile Local Model Bench is built to test. The case-by-case audit facts were often right. The final closure arithmetic was not. A model that gets the documents but trips on the checksum is useful for assisted review but not for unattended automation.
Side by side with Qwen3.6 27B on the same nine cases
Comparing Qwen3.8 27B to its predecessor on this exact suite is the most useful context. Both models ran on the same Mac mini M4 with the same LM Studio toolchain or the same Ollama runtime, weeks or months apart in the calendar. The nine cases are identical. The hidden-oracle definition is identical. The only thing that changed is the model generation and the harness.
On the strict resolved signal, Qwen3.8 27B closed 3/9 strictly: W05 email-attachment, plus none of the five generated-invoice paper-trail cases. Qwen3.6 27B closed 5/9 strictly on the same nine cases: case04 messy-intake, case06 remittance-split, case07 credit-offset, plus two of the five P-cases. Qwen3.6 hit three out of four workflow cases clean plus two P-cases; Qwen3.8 hit one out of four workflow cases clean and zero P-cases.
On the core-oracle pass, Qwen3.8 27B hit 7/9. Qwen3.6 27B hit 8/9. The lost core-oracle on Qwen3.8 is W04 messy-intake, where Qwen3.8 produced a plausible-looking but misclassified final artifact that failed on nine separate oracle checks. Qwen3.6 closed W04 clean.
Why the newer generation scored worse
On the surface this looks like a regression. The newer Qwen generation closed two fewer workflow cases strictly, and one fewer case at the core-oracle level. With nine cases and single runs, those numbers are small enough that noise plus prompt-template variance plus a different tokenizer behavior on receipts can swing them.
There are several reasonable interpretations. Qwen3.8 is documented as more agentic-friendly than Qwen3.6, with stronger tool-use and reasoning support. The agentic-optimized variant of paperwork-v3, which gives the model a tool loop of its own, scored much higher for Qwen3.8 on a separate run. The single-shot invoice cases may simply be the wrong test shape for this generation.
Another interpretation is calibration. A run of nine cases is small. Noise plus the specific prompt template plus a different tokenizer behavior on receipts can swing several percentage points. Qwen3.8 27B is real software. A single Qwen3.6 27B tier run does not prove Qwen3.6 is the better model for users. It proves Qwen3.6 is the better model for this exact benchmark on this exact date, with nine cases and one run per case.
What this means for the leaderboard
The current Local Model Bench leaderboard shows Qwen3.6 27B at the top of the local rows with a 72.2% practical score. Qwen3.8 27B entered at 55.6% on the same nine cases, slot 7 overall behind the cloud reference rows. That ranking reflects what we actually measured this week, not a final verdict on either generation.
The useful read for a private-document user is not which Qwen version is better. It is that both are good enough to take seriously, and that close generations can produce close-but-different results on the same workload. If your local stack includes Qwen3.8 27B, the assisted-review use case is fine. For unattended closure, neither generation currently qualifies.
What we would still want to know
A few things would change the picture. A second Qwen3.8 27B run on the same nine cases would show whether the 3/9 result is stable or a single-run fluctuation. The same Qwen3.8 27B on the agentic paperwork variant, where the model can call its own tools, would show whether the regression is test-shape related. A paper-trail Text-Only diagnostic, with the images stripped, would isolate how much of the failure is document-reading versus arithmetic-closure.
None of those are planned in the next session. They are the obvious follow-ups for someone who wants to understand whether Qwen3.8 27B is a real regression for private paperwork or a different shape for a different test.
Where it worked
- W05 (email-attachment intake) closed strictly across all four oracle checks. That is the strongest single workflow result in the local rows for this generation.
- Three of five generated-invoice cases reached core-oracle pass on Qwen3.8 27B. Audit-fact identification was often correct.
- No case where the model scored zero on visible-checks. The visible rule set was always at least partly honored.
- Ollama-local workflow run completed without timeouts for all four agentic cases.
Where it failed
- Zero out of five generated-invoice paper-trail cases closed strictly. Every one failed on the proof-code arithmetic at the end.
- Three of four workflow cases (W04, W06, W07) did not close the final-oracle check.
- W04 failed on nine separate oracle checks: document index, ignored documents, invoice classification, manifest, normalized text, proof code, proof file, totals, and warning codes.
- W07 failed on payment reconciliation plus several artifact files.
How the model is positioned
- Qwen's own release notes for Qwen3.8 describe it as more agentic-friendly than Qwen3.6, with stronger tool-use and reasoning support.
- The Ollama library lists qwen3.8:27b-mlx as the Qwen3.8 27B variant available for local Mac inference with metal/MLX acceleration.
- OpenRouter exposes qwen/qwen3.8-27b as a hosted endpoint on the same model card.
- Local Model Bench tests one specific practical workload, not the full model card.
What was actually tested
- The model ran all five generated-invoice Paperwork Trial cases via the OpenRouter qwen/qwen3.8-27b endpoint. Each case is a single prompt with attached invoice scans, bank exports, vendor records, and purchase orders, plus the exact audit_result.json oracle check.
- The model ran all four agentic Paperwork Workflow cases via local Ollama (qwen3.8:27b-mlx) on the Mac mini M4. Each case is a messy folder with active and draft files, a protected incoming folder, intermediate artifacts to write, and a final proof file.
- Practical Score = 0.5 × (resolved / 9) + 0.5 × (core / 9) = 0.5 × 3/9 + 0.5 × 7/9 = 55.6%.
- The Qwen3.6 27B comparison is the same nine cases on the same hardware class with the same suite version, run earlier in 2026.
Verdict
Qwen3.8 27B on paperwork-v3 is a real benchmark result: 55.6% practical, with W05 as the standout resolved workflow case. Compared against Qwen3.6 27B on the same nine cases, the newer generation closed 3/9 versus 5/9 strictly, a two-task delta inside one-task run-to-run noise on a nine-case single-run benchmark. The single-suite read is that the newer generation did not improve on this boring last mile; the more charitable read is that this exact test shape is not what Qwen3.8 was tuned for, since the agentic paperwork variant of the same suite scored higher for Qwen3.8 in a separate tool-loop run.