> ## Content Index
> Fetch the complete content index at: https://localmodelbench.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# The no-think leaderboard that wasn't: 7 rows reasoned anyway, re-measured Gemma-4 falls 61.1% to 11.1%
- URL: https://localmodelbench.com/the-no-think-leaderboard-that-wasnt/
- Published: 2026-09-25T16:51:09.000Z
- Updated: 2026-09-25T17:05:10.000Z
- Description: Seven paperwork-v3 leaderboard rows sent every no-think flag and produced reasoning tokens anyway, up to 28,821 per model. Re-measured under verified suppression, Gemma-4 26B A4B drops from 61.1% to 11.1% practical.
- Author: LMB Editorial
- Tags: paperwork-v3, methodology, gemma-4, reasoning

Seven rows on the paperwork-v3 leaderboard were measured under a no-think convention that never held. The runs sent every suppression flag we had (`reasoning.effort: none`, `think: false`, `enable_thinking: false`, a literal `/no_think` suffix in the prompt) and the models produced reasoning tokens anyway, up to 28,821 per model across its runs. When we re-ran Gemma-4 26B A4B on a runtime that actually suppresses reasoning, its practical score fell from 61.1% to 11.1%.

The audit started while preparing a thinking-enabled re-run of Qwen3.6 27B, the same exercise we did for Bonsai 2\. Its May no-think run contained multi-thousand-token reasoning traces that should not exist, so we scanned every scored run in the suite for `completion_tokens_details.reasoning_tokens` and `reasoning_content`.

## The no-think convention

Local Model Bench runs local models in a no-think configuration: the request payload sets `reasoning.effort` to `none`, `think` and `enable_thinking` to `false`, and the prompt template instructs the model not to use a reasoning trace. We assumed sending these flags made the condition real. The scan shows that assumption was wrong for the hybrid-reasoning models we probed through LM Studio's OpenAI-compatible API.

## Scan of run artefacts

Reasoning tokens per leaderboard model across the paperwork-v3 run artefacts:

| Model                  | Runs scanned | Runs with reasoning | Reasoning tokens (total across runs) |
| ---------------------- | ------------ | ------------------- | ------------------------------------ |
| prism-ml/bonsai-27b    | 5            | 5                   | 28,821                               |
| qwen/qwen3.6-35b-a3b   | 6            | 6                   | 28,095                               |
| qwen/qwen3.6-27b       | 6            | 6                   | 21,560                               |
| google/gemma-4-e4b     | 10           | 9                   | 21,299                               |
| meta/muse-glimmer      | 6            | 6                   | 16,993                               |
| google/gemma-4-26b-a4b | 5            | 4                   | 11,315                               |
| google/gemma-4-e2b     | 5            | 5                   | 9,583                                |

Three further models produced a short 299-token trace in a single run each (qwen3.5-9b, qwen3.5-35b-a3b, glm-4.7-flash), and the local nvidia/nemotron-3-nano-omni, which never completed a leaderboard row, produced 19,795 across its runs. The seven rows in the table carry the † marker described below; the 299-token models have no complete nine-case row to mark. Verified clean: prism-ml/ternary-bonsai-2-27b (its 27.8% no-think score is real, which keeps the Bonsai-2 thinking-divide comparison intact), qwen3.8-27b, gemma-4-12b, qwen3-14b, the Mistral/Ministral rows and the OpenRouter reference rows. The scan likely undercounts: the workflow runner does not persist `reasoning_content` per step, so workflow-stage reasoning is only visible where usage accounting exposes it.

## Gemma-4 26B re-measure

We re-ran Gemma-4 26B A4B, one of the flagged rows, on the same nine cases through llama.cpp with `--reasoning off`, the only suppression path we found that produces zero reasoning tokens. Practical is the leaderboard score: 0.5 × strict + 0.5 × core.

| Configuration                         | Strict | Core | Practical | Reasoning tokens |
| ------------------------------------- | ------ | ---- | --------- | ---------------- |
| LM Studio MLX, "no-think" flags (May) | 4/9    | 7/9  | 61.1%     | 11,315 hidden    |
| llama.cpp \--reasoning off (verified) | 0/9    | 2/9  | 11.1%     | 0                |

Practical fell 50.0 points, larger than the 44.4-point thinking divide we published for Bonsai 2\. The May row included 11,315 hidden reasoning tokens.

## Failure signature

All five generated-invoice cases failed the core oracle on the same field: `proof_code`, a deterministic checksum over approved amounts, invoice-ID numerics and warning counts. Two cases make the mechanism unusually visible:

- **P01 and P02: perfect audit, wrong arithmetic.** In P01 the no-think output approved INV-7801, flagged INV-7802 (payment\_short) and INV-8422 (under\_review\_stamp), ignored the QT-6400 quote and totalled 18,737 gross cents: identical to the oracle on every classification field. Only the proof code was wrong, 20,104 instead of 42,956\. The model read every document correctly and lost the case on multi-digit mental arithmetic it could not offload to a reasoning trace.
- **P03–P05: subtle verdict drift.** Each case approved exactly one invoice the oracle wanted held for review (INV-7801, INV-4171, INV-5601 respectively), alongside the same broken proof codes.

The signature differs from Bonsai 2's no-think failure. Bonsai 2 emitted the proof code as an unevaluated expression and broke JSON validity; Gemma-4 evaluated its formula and returned a wrong number that still parses.

The verified no-think generated cases were also much faster: 7–18 seconds per case by run-directory timestamps, versus 0.5–2 minutes per case in the May MLX runs.

## Suppression probes

We probed six suppression paths through LM Studio's API on both Qwen3.6 27B and Gemma-4 26B: baseline flags, the original May configuration, `chat_template_kwargs.enable_thinking`, a bare `/no_think` suffix, and combinations. Every call produced reasoning content (299 to 3,529 tokens) regardless of which flags were set. LM Studio applies the flags to the request, but these chat templates do not honour them; the model decides.

The only suppression that held was server-side: llama.cpp's `--reasoning off` (we used the PrismML fork at b10735, needed anyway for the gemma4uv projector type). With it, probes returned `reasoning_content` empty and zero reasoning tokens.

## Leaderboard changes

- Every row now carries a **Mode** label. The seven rows in the scan table are labelled **thinking** (that is what they were) and keep their **†** marker with the token count on hover. The three 299-token single-run models have no complete nine-case row, so nothing is marked for them.
- Gemma-4 26B A4B gets a second row, `gemma-4-26b-a4b-nothink` at 11.1%: the first leaderboard entry measured under verified reasoning suppression (reasoning\_tokens = 0 in every response).
- The default convention is now **thinking-enabled**: the leaderboard's primary score measures what a model can do with its reasoning enabled, rather than what it can be forced not to do. The seven flagged rows keep their numbers under the thinking label; they were honestly measured, just mislabelled.
- Verified **no-think** rows keep their scores as the explicit secondary condition and are being re-measured under thinking over time. As a label, "no-think" now means the recorded response contains zero reasoning tokens, not that the request sent suppression flags.

## Caveats

- **Engine confound.** The verified no-think run moved from LM Studio's MLX build to llama.cpp with a Q4\_K\_XL quant, so two variables changed, not one. The failure pattern argues for a reasoning effect rather than a quantisation effect (correct classification plus wrong arithmetic is a trace dependency, not a quality loss), but a same-engine A/B would be cleaner and is only possible where a runtime can suppress reasoning at all.
- **Single runs.** Both rows are one pass per case. On nine cases a single case swinging between pass and fail moves the practical score by 5.6 to 11.1 points; repeated runs of the same configuration have differed by about that much. The 50-point movement here is more than four times that.
- **Incomplete telemetry.** Models whose reasoning happened only inside workflow steps may not appear in the scan. The marked set is a floor, not a ceiling.
- The Bonsai-2 note states every other leaderboard model ran under the no-think convention. Technically true (the flags were sent), but we have added a footnote there pointing here.

## Method and artefacts

- Scan: every `raw_response.json` under `results/paperwork_v3_case_runs/` and the workflow run roots, for `reasoning_tokens` in usage and non-empty `reasoning_content`. Artefact: `tmp/hidden_reasoning_scan.json`.
- Suppression probes: six payload variants per model through LM Studio's API; llama.cpp `--reasoning off` on Qwen3.6-27B Q4\_K\_M, Gemma-4-12B Q4\_0 and Gemma-4-26B-A4B QAT Q4\_K\_XL.
- Re-measurement: all nine paperwork-v3 cases, unsloth `gemma-4-26B-A4B-it-qat-Q4_K_XL.gguf` \+ BF16 mmproj, PrismML llama.cpp fork b10735, `-ngl 99 -fa on -c 65536 --reasoning off`, temperature 0, byte-identical prompts to the May run.
- Score check: 0.5×4/9 + 0.5×7/9 = 61.1%; 0.5×0/9 + 0.5×2/9 = 11.1%.

## Verdict

> Seven leaderboard rows were measured under a no-think convention the models never honoured: sending the flags did not stop the reasoning trace. Re-measured under verified suppression, Gemma-4 26B A4B's practical score fell from 61.1% to 11.1%, subject to the engine confound noted above. The leaderboard now tests thinking-enabled by default; no-think remains as a secondary condition, and it means zero reasoning tokens in the recorded response, not flags in the request.

## Sources

- [unsloth/gemma-4-26B-A4B-it-qat-GGUF (Q4\_K\_XL + mmproj-BF16)](https://huggingface.co/unsloth/gemma-4-26B-A4B-it-qat-GGUF?ref=localmodelbench.com)
- [PrismML llama.cpp fork, with \--reasoning off and gemma4uv projector support](https://github.com/PrismML-Eng/llama.cpp?ref=localmodelbench.com)
- [Bonsai 2 27B note, the thinking-divide measurement this audit builds on](https://localmodelbench.com/bonsai-2-27b-on-paperwork-v3-thinking-divide/)
- [Methodology: how strict, core and practical are scored](https://localmodelbench.com/methodology/)