> ## Content Index
> Fetch the complete content index at: https://localmodelbench.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Reasoning budget beats raw size on local math
- URL: https://localmodelbench.com/reasoning-budget-beats-raw-size/
- Published: 2026-09-15T19:16:49.000Z
- Updated: 2026-09-15T19:16:49.000Z
- Description: A 27B dense model with reasoning on solves all 36 math tasks. Reasoning budget beats raw size on local math. Tested 8 model configs across 9 stress dimensions.
- Author: Local Model Bench Admin
- Tags: field-note

*Field note — Local Model Bench — September 2026*

## Short answer

A 27B dense model with reasoning on solves all 36 math tasks in this probe. The 27B dense model one generation newer, with reasoning on, solves 35\. A 26B MoE with reasoning on (only 4B active parameters) also solves 35 of 36 — only the 12-digit sum times out. Switch reasoning off on those same models and the dense 27B drops to 28 of 36, the 4B-active MoE drops to 31\. Drop the size and an 8B solves 20, a 14B solves 24.

Size helps. Reasoning helps more. They interact, and they only get you to 100% when both are stacked on the right architecture — and the right configuration matters as much as the right model. Generation does not help uniformly: the newer Qwen loses ground on 12-digit addition that the older Qwen solved.

## What the probe measures

36 tasks across 9 stress dimensions, four tasks per dimension. Original numbers throughout — no GSM8K, no MATH, no public benchmark contamination. Tolerance is exact or ±0.01 / ±0.1 / ±1.0 depending on the task.

| Dimension           | What scales                                    |
| ------------------- | ---------------------------------------------- |
| Digit count         | digits per operand (2 → 4 → 6 → 12)            |
| Operand count       | operands to sum (2 → 5 → 10 → 20)              |
| Chain depth         | sequential operations (1 → 2 → 3 → 5 steps)    |
| Decimal precision   | cents → 4-decimal sum                          |
| Signed numbers      | clean → 1 negative → mixed signs → large mixed |
| Word problems       | with irrelevant numbers in the prompt          |
| Unit conversion     | as part of the arithmetic                      |
| Named variables     | dependency chain (1 → 2 → 3 → 5 steps)         |
| Multiplication size | 1×1 → 5×5 digits                               |

## Models and configurations

| Model            | Size            | Architecture | Reasoning | Score              |
| ---------------- | --------------- | ------------ | --------- | ------------------ |
| Granite 4.1 8B   | 8B              | dense        | no        | 20/36 (55.6%)      |
| Phi-4            | 14B             | dense        | no        | 24/36 (66.7%)      |
| Qwen 3.6 27B     | 27B             | dense        | off       | 28/36 (77.8%)      |
| Qwen 3.8 27B     | 27B             | dense        | off       | 27/36 (75.0%)      |
| Gemma 4 26B-A4B  | 26B (4B active) | MoE          | off       | 31/36 (86.1%)      |
| Gemma 4 26B-A4B  | 26B (4B active) | MoE          | on        | 35/36 (97.2%)      |
| Qwen 3.8 27B     | 27B             | dense        | on        | 35/36 (97.2%)      |
| **Qwen 3.6 27B** | 27B             | dense        | on        | **36/36 (100.0%)** |

`temperature=0`, `top_p=1`. Reasoning-on runs used `max_tokens=8192`; the rest used `max_tokens=4096`. Same system prompt for all: "you are a careful arithmetic calculator, give only the final number."

## Per-dimension accuracy

| Dimension         | Granite 8B | Phi-4 14B | Qwen 3.6 off | Qwen 3.8 off | Gemma MoE off | Gemma MoE on | Qwen 3.8 on | Qwen 3.6 on |
| ----------------- | ---------- | --------- | ------------ | ------------ | ------------- | ------------ | ----------- | ----------- |
| Digit count       | 2/4        | 3/4       | 4/4          | 3/4          | 3/4           | 3/4          | 3/4         | **4/4**     |
| Operand count     | 3/4        | 3/4       | 3/4          | 4/4          | 4/4           | **4/4**      | **4/4**     | **4/4**     |
| Chain depth       | 1/4        | 2/4       | 2/4          | 2/4          | 3/4           | **4/4**      | **4/4**     | **4/4**     |
| Decimal precision | 4/4        | 4/4       | 4/4          | 4/4          | 4/4           | 4/4          | 4/4         | 4/4         |
| Signed numbers    | 2/4        | 3/4       | 3/4          | 2/4          | 3/4           | **4/4**      | **4/4**     | **4/4**     |
| Word problems     | 4/4        | 4/4       | 4/4          | 4/4          | 4/4           | 4/4          | 4/4         | 4/4         |
| Unit conversion   | 1/4        | 2/4       | 3/4          | 3/4          | 4/4           | **4/4**      | **4/4**     | **4/4**     |
| Named variables   | 1/4        | 1/4       | 2/4          | 2/4          | 3/4           | **4/4**      | **4/4**     | **4/4**     |
| Multiplication    | 2/4        | 2/4       | 3/4          | 3/4          | 3/4           | **4/4**      | **4/4**     | **4/4**     |

Every model is perfect on decimal precision and word problems. On multiplication, the two generation-different Qwens tie at 3/4 with reasoning off; all three reasoning-on configurations solve it cleanly.

## The same model, two configurations (Qwen 3.6)

The cleanest signal in the probe is comparing Qwen 3.6 27B to itself:

- **Reasoning off** (max\_tokens=4096): 28/36 (77.8%), 1.47 s per task.
- **Reasoning on** (max\_tokens=8192): 36/36 (100.0%), \~38 s per task.

Same weights, same temperature, same system prompt. The only difference is whether the model uses its internal scratchpad channel before emitting a final answer. The 26× slowdown is real.

## Generation drift (Qwen 3.6 vs Qwen 3.8)

The two 27B dense models differ on two tasks:

- **A4 — 12-digit + 12-digit + 12-digit.** Qwen 3.6 ON solves it exactly. Qwen 3.8 ON misses by 300 thousand (1585059803977 vs 1585059503977). Same in the reasoning-off configuration: 3.6 solves it exactly, 3.8 misses by 40 thousand (1585059464077).
- **I4 — 84736 × 59382.** Qwen 3.8 ON solves it exactly. Qwen 3.6 ON also solves it exactly. In the reasoning-off configuration, Qwen 3.6 is off by 62 thousand and Qwen 3.8 is off by 92 thousand — both within 0.002% of the true product, both close, and 3.6 marginally closer.

## What still breaks with reasoning on

After reasoning is on, two models have exactly one wall left each:

- **A4 — Gemma 4 26B-A4B.** Times out at 200 s in the main probe and at 409 s in a bonus run with `max_tokens=16384` and a 600-second timeout. Only 4B active parameters at inference; 12-digit chain-of-thought does not finish inside a usable budget.
- **A4 — Qwen 3.8 27B.** Returns a wrong answer (off by 300 thousand on a 12-digit number) rather than timing out. With reasoning on and 8192 token budget, the model emits confidently and incorrectly.

Qwen 3.6 ON crosses every wall. Gemma-on crosses every wall except A4\. Qwen 3.8 ON crosses every wall except A4.

## What still breaks at smaller sizes (or without reasoning)

- **Chain depth (3+ steps).** Granite 1/4, Phi-4 2/4, all three no-reasoning 27B/26B-MoE configs 2-3/4\. Both reasoning-on dense configs 4/4.
- **Named-variable chains.** Same shape: Granite 1/4, Phi-4 1/4, no-reasoning 2-3/4, reasoning-on 4/4.
- **Large mixed-sign sums.** Granite, Phi-4, Gemma-off all miss. Both Qwen-on configs and Gemma-on 4/4.
- **Mixed-unit arithmetic.** Granite 1/4, Phi-4 2/4, Qwen-off 3/4\. Reasoning-on configs and Gemma-off 4/4.
- **12-digit +.** Granite, Phi-4, Gemma-off, and both Qwen 3.8 configs miss. Qwen 3.6 off and Qwen 3.6 on solve it. Gemma-on times out.

Below reasoning-on-Qwen 3.6, every one of these is a "verify or use a tool" case. Reasoning on a large dense model is what crosses them.

## The MoE surprise

Gemma 4 26B-A4B has 26B total parameters but only 4B active at inference. With reasoning off, it scores 31/36 — beating both 27B dense Qwens (Qwen 3.6 28, Qwen 3.8 27) at sub-second latency. With reasoning on, it lands at 35/36 (97.2%) at about 8 seconds per task. The MoE reasoning latency is much lower than the dense 27B reasoning latency measured via LM Studio (\~40 s for Qwen 3.6; Gemma-on at \~15 s per task). Qwen 3.8 via Ollama comes in at \~6 s per task. Most of that gap is LM Studio vs Ollama, not the model. Reasoning-on dense remains slow.

The MoE wins on reasoning-off latency and accuracy. The dense 27B Qwen 3.6 wins on the last hard items. A4 is the 4B-active ceiling.

## Practical readout

If you are running local models and you need them to do math:

- **Decimals, percent, currency, word problems:** any model. Don't burn reasoning budget here.
- **Two- to four-digit arithmetic:** any model.
- **Mixed-unit arithmetic (cm + mm, kg + g):** Gemma 4 26B-A4B or Qwen 3.6 / 3.8 27B with reasoning on. Phi-4 and Granite 4.1 8B fail at this.
- **Multi-step chains with named variables, 3+ steps:** the three reasoning-on configurations handle all four cases. Below that, verify.
- **Large mixed-sign sums (7+ operands):** the three reasoning-on configurations solved this. Below that, verify or use a tool.
- **12-digit addition:** Qwen 3.6 ON solves it exactly. Qwen 3.8 ON misses by 300 thousand. Gemma-on times out. Below that, verify.
- **Multiplication past 3 digits:** all three reasoning-on configurations solved it exactly (5,031,793,793). Below that, verify.

If you are deciding between 8B / 14B / 27B and care about math:

- The 8B Granite is fast and accurate on the easy half.
- The 14B Phi-4 buys \~11 percentage points over the 8B on the full probe.
- The 27B Qwen 3.6 with reasoning on solves the entire probe. Adds \~40 s of latency per task (LM Studio backend). The right choice when accuracy matters and latency is fine.
- The 27B Qwen 3.8 with reasoning on solves 35/36 at \~6 s per task (Ollama backend). The latency gap to Qwen 3.6 is mostly the backend, not the model. At the cost of 12-digit accuracy.
- The 26B MoE Gemma without reasoning lands between the 14B and the reasoning-on 27B on score, at sub-second latency. The math-strongest non-reasoning option in the probe.
- The 26B MoE Gemma with reasoning on lands at 35/36 (97.2%) at \~15 s per task (LM Studio backend). The 27B Qwen 3.8 with reasoning on also lands at 35/36 (97.2%) at \~6 s per task (Ollama backend). Pick by backend: under LM Studio, Gemma-on is the fastest 35/36; under Ollama, Qwen 3.8-on is.

If your pipeline calls the model hundreds of times in a loop, the reasoning-on 27B is not the right choice. The reasoning-off 26B-A4B or the reasoning-off 27B Qwen is.

If your pipeline calls the model once per document and accuracy matters, the reasoning-on Qwen 3.6 is the right choice.

## Probe details

- 36 tasks, all original numbers
- `temperature=0`, `top_p=1`
- Reasoning-off and non-reasoning runs: `max_tokens=4096`
- Reasoning-on runs: `max_tokens=8192`
- Same system prompt: "you are a careful arithmetic calculator, give only the final number"
- LM Studio OpenAI-compatible endpoint at `http://localhost:1234/v1` for Qwen 3.6 and Gemma
- Ollama OpenAI-compatible endpoint at `http://localhost:11434/v1` for Granite 4.1 8B and Qwen 3.8
- Reasoning models with default LM Studio settings burn their entire token budget on the internal scratchpad without emitting a final answer. For automated extraction pipelines, set `reasoning_effort: "none"` to disable the scratchpad, or raise `max_tokens` to 8192+ to give it room to complete.
- Raw responses + grades: `tmp/math-probe/results-*.json`
- Runner + analyzer: `tmp/math-probe/run_probe.py`, `analyze.py`

## Bottom line

Size helps on the easy half of the math probe — 8B → 14B → 27B moves 20 → 24 → 28 with reasoning off. Reasoning on top of size crosses the rest: Qwen 3.6 27B with reasoning on jumps from 28 to 36\. The MoE with only 4B active parameters comes within one task of the dense 27B at a quarter of the reasoning latency. The newer Qwen 3.8 is one task behind. Pick by backend and latency budget.