> ## Content Index
> Fetch the complete content index at: https://localmodelbench.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Muse Glimmer: 6/9 on paperwork-v3. The first run returned empty content.
- URL: https://localmodelbench.com/muse-glimmer-paperwork-v3/
- Published: 2026-09-18T16:40:34.000Z
- Updated: 2026-09-23T17:06:03.000Z
- Description: Muse-Glimmer scored 6/9 strict and 8/9 core on paperwork-v3 (77.8% practical), ranking as the top local model. Its first run returned empty output after 2,397 reasoning tokens exhausted the default 2,400-token limit. Raising max_tokens to 8,000 resolved the bottleneck.
- Author: LMB Editorial

meta/muse-glimmer is a 30B reasoning model running locally in LM Studio on Apple Silicon. On the nine-case paperwork-v3 suite, it scored 6/9 strict and 8/9 core, for a practical\_score of 77.8%. That ties it with OpenAI GPT-5.4 Mini (Codex CLI) at 77.8%. On the posted order it sits fourth; it is the highest local row, ahead of qwen3.6-27b at 72.2%. The initial smoke run on p01 returned an empty response because reasoning tokens exhausted the completion limit; raising the token ceiling produced a completed JSON body.

## What ran

Muse-Glimmer ran inside LM Studio on a Mac mini M4 (64 GB unified memory) and was queried via the local OpenAI-compatible endpoint at `http://localhost:1234/v1`. Sampling was set to `temperature: 0` and `top_p: 1` (Meta's defaults are `temperature: 1.0`, `top_p: 0.95`, `top_k: 64`). The benchmark evaluated five generated-invoice cases (p01–p05) and four multi-step workflow cases (w04–w07).

| Case                                      | Type     | core\_ok | hidden\_oracle     | elapsed | completion tokens | of which reasoning |
| ----------------------------------------- | -------- | -------- | ------------------ | ------- | ----------------- | ------------------ |
| p01 (generated\_invoice\_case\_01)        | image    | yes      | yes                | 3:26    | 2 681             | 2 468              |
| p02 (generated\_invoice\_case\_02)        | image    | yes      | yes                | 3:04    | 1 852             | 1 690              |
| p03 (generated\_invoice\_case\_03)        | image    | no       | no                 | 9:20    | 6 099             | 5 758              |
| p04 (generated\_invoice\_case\_04)        | image    | yes      | yes                | 4:03    | 2 632             | 2 489              |
| p05 (generated\_invoice\_case\_05)        | image    | yes      | yes                | 3:37    | 2 303             | 2 191              |
| w04 (messy\_intake\_workflow\_case\_04)   | workflow | yes      | no (manifest)      | 6:00    | 4 147             | n/a                |
| w05 (email\_attachment\_intake\_case\_05) | workflow | yes      | no (document\_set) | 10:00   | \~7 800           | n/a                |
| w06 (remittance\_split\_case\_06)         | workflow | yes      | yes                | 7:22    | \~7 900           | n/a                |
| w07 (credit\_offset\_case\_07)            | workflow | yes      | yes                | 9:11    | \~9 200           | n/a                |

Six clean hidden-oracle passes (p01, p02, p04, p05, w06, w07), two near misses where core audit values matched but artifact files failed schema verification (w04, w05), and one failure (p03). This tallies to 6/9 strict (66.7%) and 8/9 core (88.9%). The practical\_score equals 0.5 × (6/9) + 0.5 × (8/9) = 14/18 = 77.8%.

## The first run returned empty content

The initial run on case\_01 used `max_tokens: 2400`, the standard ceiling applied to non-reasoning runs in this harness. The model generated 2 397 tokens of hidden reasoning, hit the limit, and returned an empty string in `content` with `finish_reason: length`.

Standard OpenAI-compatible parameters designed to suppress reasoning (`think: false`, `enable_thinking: false`, `reasoning: {effort: "none"}`) had no effect in the request body. The Jinja chat template bundled with llama.cpp inside LM Studio did not evaluate those fields. The only operational control exposed by the runner was the token ceiling.

Raising `max_tokens` to 8 000 resolved the failure. With an 8 000-token limit, the reasoning trace on p01 used 2 468 tokens, followed by a valid 213-token JSON object containing all required audit fields.

| max\_tokens | reasoning\_tokens | completion\_tokens | content   | finish\_reason | core\_ok |
| ----------- | ----------------- | ------------------ | --------- | -------------- | -------- |
| 2 400       | 2 397             | 2 400              | ""        | length         | no       |
| 8 000       | 2 468             | 2 681              | full JSON | stop           | yes      |

An alternative method is setting `chat_template_kwargs: {"enable_thinking": false}` in the request body. In this harness, llama.cpp parsed that parameter for the Muse-Glimmer template, removing the reasoning trace entirely. On w04, disabling thinking reduced run time from 6:00 to 5:30 while maintaining the same `core_ok = true` result and identical schema validation errors.

## What p03 missed

Across the five invoice cases, four passed all visible and hidden checks. The sole failure occurred on p03 due to an evidence-handling error rather than a token constraint.

Case p03 provides a `previous_invoices.csv` ledger in the case folder. The reference oracle flags invoice `INV-7801` with the warning code `duplicate_risk` because the record appeared in an earlier billing cycle. Muse-Glimmer spent 5 758 reasoning tokens analyzing the files but approved `INV-7801` instead of routing it to review. This doubled `total_approved_gross_cents` from 18 737 to 37 474 and corrupted the checksum in `proof_code`.

Re-running p03 confirmed this was not a generation limit issue. The model parsed the scans and CSVs correctly, but failed to apply the duplicate-detection rule defined in the prompt.

## What w04 and w05 missed

The four workflow cases (w04–w07) test an agentic tool loop: inspecting files, writing intermediate files, and producing final deliverables.

On w04 (intake of mixed scans and stale exports), Muse-Glimmer identified the correct audit values (`core_ok: true`), but generated `work/normalized_manifest.json` as a flat JSON dictionary without the required top-level `normalized_manifest` wrapper key. The local evaluator marked the artifact as a structural failure (`manifest_error`).

A similar issue occurred on w05: the core audit fields matched ground truth, but `work/final_document_set.json` recorded two documents instead of the three expected by the hidden oracle (`final_document_set_error`). Cases w06 and w07 completed all steps without errors, passing both core and hidden oracles.

## Where this puts the leaderboard

The paperwork-v3 leaderboard computes `practical_score` as an equal blend of strict passes (100% hidden-oracle compliance) and core passes (accurate financial values regardless of file wrapper defects):

| Rank  | Model                           | practical\_score | strict pass | core pass |
| ----- | ------------------------------- | ---------------- | ----------- | --------- |
| 1     | opencode/minimax-m3-free        | 88.9%            | 8/9         | 8/9       |
| 2     | OpenAI GPT-5.5 (Codex CLI)      | 83.3%            | 7/9         | 8/9       |
| 3     | OpenAI GPT-5.4 Mini (Codex CLI) | 77.8%            | 7/9         | 7/9       |
| **4** | **meta/muse-glimmer**           | **77.8%**        | **6/9**     | **8/9**   |
| 5     | qwen3.6-27b                     | 72.2%            | 5/9         | 8/9       |
| 6     | antigravity-gemini-3.8-flash    | 72.2%            | 4/9         | 9/9       |
| 7     | google/gemma-4-26b-a4b          | 61.1%            | 4/9         | 7/9       |
| 8     | qwen3.8-27b                     | 55.6%            | 3/9         | 7/9       |
| 9     | qwen3.6-35b-a3b                 | 38.9%            | 1/9         | 6/9       |

Muse-Glimmer ties GPT-5.4 Mini at 77.8% and leads all locally run open-weights models on this hardware. The gap over qwen3.6-27b (72.2%) is one extra strict case (6/9 vs 5/9) at the same 8/9 core rate.

## Harness requirements for reasoning models

Evaluating reasoning models in local harnesses requires sizing token ceilings for thinking traces. The default 2 400-token ceiling designed for direct-instruct models caused an immediate generation failure on p01\. An 8 000-token ceiling left room for the reasoning trace and a finished JSON body.

For reproducible local benchmarking, runners need either configurable per-model token limits or direct exposure of template controls such as `chat_template_kwargs`. Without these adjustments, standard runner presets test token truncation limits rather than model reasoning capability.