Muse Glimmer: 6/9 on paperwork-v3. The first run returned empty content.

Muse-Glimmer scored 6/9 strict and 8/9 core on paperwork-v3 (77.8% practical), ranking as the top local model. Its first run returned empty output after 2,397 reasoning tokens exhausted the default 2,400-token limit. Raising max_tokens to 8,000 resolved the bottleneck.

Thermal-printer receipt: INVOICE-AUDIT-RUN / content: "" / finish_reason: length, rest blank. Documentaristic still-life in the LMB paper-line.

meta/muse-glimmer is a 30B reasoning model running locally in LM Studio on Apple Silicon. On the nine-case paperwork-v3 suite, it scored 6/9 strict and 8/9 core, for a practical_score of 77.8%. That ties it with OpenAI GPT-5.4 Mini (Codex CLI) at 77.8%. On the posted order it sits fourth; it is the highest local row, ahead of qwen3.6-27b at 72.2%. The initial smoke run on p01 returned an empty response because reasoning tokens exhausted the completion limit; raising the token ceiling produced a completed JSON body.

What ran

Muse-Glimmer ran inside LM Studio on a Mac mini M4 (64 GB unified memory) and was queried via the local OpenAI-compatible endpoint at http://localhost:1234/v1. Sampling was set to temperature: 0 and top_p: 1 (Meta's defaults are temperature: 1.0, top_p: 0.95, top_k: 64). The benchmark evaluated five generated-invoice cases (p01–p05) and four multi-step workflow cases (w04–w07).

CaseTypecore_okhidden_oracleelapsedcompletion tokensof which reasoning
p01 (generated_invoice_case_01)imageyesyes3:262 6812 468
p02 (generated_invoice_case_02)imageyesyes3:041 8521 690
p03 (generated_invoice_case_03)imagenono9:206 0995 758
p04 (generated_invoice_case_04)imageyesyes4:032 6322 489
p05 (generated_invoice_case_05)imageyesyes3:372 3032 191
w04 (messy_intake_workflow_case_04)workflowyesno (manifest)6:004 147n/a
w05 (email_attachment_intake_case_05)workflowyesno (document_set)10:00~7 800n/a
w06 (remittance_split_case_06)workflowyesyes7:22~7 900n/a
w07 (credit_offset_case_07)workflowyesyes9:11~9 200n/a

Six clean hidden-oracle passes (p01, p02, p04, p05, w06, w07), two near misses where core audit values matched but artifact files failed schema verification (w04, w05), and one failure (p03). This tallies to 6/9 strict (66.7%) and 8/9 core (88.9%). The practical_score equals 0.5 × (6/9) + 0.5 × (8/9) = 14/18 = 77.8%.

The first run returned empty content

The initial run on case_01 used max_tokens: 2400, the standard ceiling applied to non-reasoning runs in this harness. The model generated 2 397 tokens of hidden reasoning, hit the limit, and returned an empty string in content with finish_reason: length.

Standard OpenAI-compatible parameters designed to suppress reasoning (think: false, enable_thinking: false, reasoning: {effort: "none"}) had no effect in the request body. The Jinja chat template bundled with llama.cpp inside LM Studio did not evaluate those fields. The only operational control exposed by the runner was the token ceiling.

Raising max_tokens to 8 000 resolved the failure. With an 8 000-token limit, the reasoning trace on p01 used 2 468 tokens, followed by a valid 213-token JSON object containing all required audit fields.

max_tokensreasoning_tokenscompletion_tokenscontentfinish_reasoncore_ok
2 4002 3972 400""lengthno
8 0002 4682 681full JSONstopyes

An alternative method is setting chat_template_kwargs: {"enable_thinking": false} in the request body. In this harness, llama.cpp parsed that parameter for the Muse-Glimmer template, removing the reasoning trace entirely. On w04, disabling thinking reduced run time from 6:00 to 5:30 while maintaining the same core_ok = true result and identical schema validation errors.

What p03 missed

Across the five invoice cases, four passed all visible and hidden checks. The sole failure occurred on p03 due to an evidence-handling error rather than a token constraint.

Case p03 provides a previous_invoices.csv ledger in the case folder. The reference oracle flags invoice INV-7801 with the warning code duplicate_risk because the record appeared in an earlier billing cycle. Muse-Glimmer spent 5 758 reasoning tokens analyzing the files but approved INV-7801 instead of routing it to review. This doubled total_approved_gross_cents from 18 737 to 37 474 and corrupted the checksum in proof_code.

Re-running p03 confirmed this was not a generation limit issue. The model parsed the scans and CSVs correctly, but failed to apply the duplicate-detection rule defined in the prompt.

What w04 and w05 missed

The four workflow cases (w04–w07) test an agentic tool loop: inspecting files, writing intermediate files, and producing final deliverables.

On w04 (intake of mixed scans and stale exports), Muse-Glimmer identified the correct audit values (core_ok: true), but generated work/normalized_manifest.json as a flat JSON dictionary without the required top-level normalized_manifest wrapper key. The local evaluator marked the artifact as a structural failure (manifest_error).

A similar issue occurred on w05: the core audit fields matched ground truth, but work/final_document_set.json recorded two documents instead of the three expected by the hidden oracle (final_document_set_error). Cases w06 and w07 completed all steps without errors, passing both core and hidden oracles.

Where this puts the leaderboard

The paperwork-v3 leaderboard computes practical_score as an equal blend of strict passes (100% hidden-oracle compliance) and core passes (accurate financial values regardless of file wrapper defects):

RankModelpractical_scorestrict passcore pass
1opencode/minimax-m3-free88.9%8/98/9
2OpenAI GPT-5.5 (Codex CLI)83.3%7/98/9
3OpenAI GPT-5.4 Mini (Codex CLI)77.8%7/97/9
4meta/muse-glimmer77.8%6/98/9
5qwen3.6-27b72.2%5/98/9
6antigravity-gemini-3.8-flash72.2%4/99/9
7google/gemma-4-26b-a4b61.1%4/97/9
8qwen3.8-27b55.6%3/97/9
9qwen3.6-35b-a3b38.9%1/96/9

Muse-Glimmer ties GPT-5.4 Mini at 77.8% and leads all locally run open-weights models on this hardware. The gap over qwen3.6-27b (72.2%) is one extra strict case (6/9 vs 5/9) at the same 8/9 core rate.

Harness requirements for reasoning models

Evaluating reasoning models in local harnesses requires sizing token ceilings for thinking traces. The default 2 400-token ceiling designed for direct-instruct models caused an immediate generation failure on p01. An 8 000-token ceiling left room for the reasoning trace and a finished JSON body.

For reproducible local benchmarking, runners need either configurable per-model token limits or direct exposure of template controls such as chat_template_kwargs. Without these adjustments, standard runner presets test token truncation limits rather than model reasoning capability.