Reasoning budget beats raw size on local math

A 27B dense model with reasoning on solves all 36 math tasks. Reasoning budget beats raw size on local math. Tested 8 model configs across 9 stress dimensions.

Reasoning budget beats raw size on local math

Field note — Local Model Bench — September 2026

Short answer

A 27B dense model with reasoning on solves all 36 math tasks in this probe. The 27B dense model one generation newer, with reasoning on, solves 35. A 26B MoE with reasoning on (only 4B active parameters) also solves 35 of 36 — only the 12-digit sum times out. Switch reasoning off on those same models and the dense 27B drops to 28 of 36, the 4B-active MoE drops to 31. Drop the size and an 8B solves 20, a 14B solves 24.

Size helps. Reasoning helps more. They interact, and they only get you to 100% when both are stacked on the right architecture — and the right configuration matters as much as the right model. Generation does not help uniformly: the newer Qwen loses ground on 12-digit addition that the older Qwen solved.

What the probe measures

36 tasks across 9 stress dimensions, four tasks per dimension. Original numbers throughout — no GSM8K, no MATH, no public benchmark contamination. Tolerance is exact or ±0.01 / ±0.1 / ±1.0 depending on the task.

Dimension What scales
Digit count digits per operand (2 → 4 → 6 → 12)
Operand count operands to sum (2 → 5 → 10 → 20)
Chain depth sequential operations (1 → 2 → 3 → 5 steps)
Decimal precision cents → 4-decimal sum
Signed numbers clean → 1 negative → mixed signs → large mixed
Word problems with irrelevant numbers in the prompt
Unit conversion as part of the arithmetic
Named variables dependency chain (1 → 2 → 3 → 5 steps)
Multiplication size 1×1 → 5×5 digits

Models and configurations

Model Size Architecture Reasoning Score
Granite 4.1 8B 8B dense no 20/36 (55.6%)
Phi-4 14B dense no 24/36 (66.7%)
Qwen 3.6 27B 27B dense off 28/36 (77.8%)
Qwen 3.8 27B 27B dense off 27/36 (75.0%)
Gemma 4 26B-A4B 26B (4B active) MoE off 31/36 (86.1%)
Gemma 4 26B-A4B 26B (4B active) MoE on 35/36 (97.2%)
Qwen 3.8 27B 27B dense on 35/36 (97.2%)
Qwen 3.6 27B 27B dense on 36/36 (100.0%)

temperature=0, top_p=1. Reasoning-on runs used max_tokens=8192; the rest used max_tokens=4096. Same system prompt for all: "you are a careful arithmetic calculator, give only the final number."

Per-dimension accuracy

Dimension Granite 8B Phi-4 14B Qwen 3.6 off Qwen 3.8 off Gemma MoE off Gemma MoE on Qwen 3.8 on Qwen 3.6 on
Digit count 2/4 3/4 4/4 3/4 3/4 3/4 3/4 4/4
Operand count 3/4 3/4 3/4 4/4 4/4 4/4 4/4 4/4
Chain depth 1/4 2/4 2/4 2/4 3/4 4/4 4/4 4/4
Decimal precision 4/4 4/4 4/4 4/4 4/4 4/4 4/4 4/4
Signed numbers 2/4 3/4 3/4 2/4 3/4 4/4 4/4 4/4
Word problems 4/4 4/4 4/4 4/4 4/4 4/4 4/4 4/4
Unit conversion 1/4 2/4 3/4 3/4 4/4 4/4 4/4 4/4
Named variables 1/4 1/4 2/4 2/4 3/4 4/4 4/4 4/4
Multiplication 2/4 2/4 3/4 3/4 3/4 4/4 4/4 4/4

Every model is perfect on decimal precision and word problems. On multiplication, the two generation-different Qwens tie at 3/4 with reasoning off; all three reasoning-on configurations solve it cleanly.

The same model, two configurations (Qwen 3.6)

The cleanest signal in the probe is comparing Qwen 3.6 27B to itself:

  • Reasoning off (max_tokens=4096): 28/36 (77.8%), 1.47 s per task.
  • Reasoning on (max_tokens=8192): 36/36 (100.0%), ~38 s per task.

Same weights, same temperature, same system prompt. The only difference is whether the model uses its internal scratchpad channel before emitting a final answer. The 26× slowdown is real.

Generation drift (Qwen 3.6 vs Qwen 3.8)

The two 27B dense models differ on two tasks:

  • A4 — 12-digit + 12-digit + 12-digit. Qwen 3.6 ON solves it exactly. Qwen 3.8 ON misses by 300 thousand (1585059803977 vs 1585059503977). Same in the reasoning-off configuration: 3.6 solves it exactly, 3.8 misses by 40 thousand (1585059464077).
  • I4 — 84736 × 59382. Qwen 3.8 ON solves it exactly. Qwen 3.6 ON also solves it exactly. In the reasoning-off configuration, Qwen 3.6 is off by 62 thousand and Qwen 3.8 is off by 92 thousand — both within 0.002% of the true product, both close, and 3.6 marginally closer.

What still breaks with reasoning on

After reasoning is on, two models have exactly one wall left each:

  • A4 — Gemma 4 26B-A4B. Times out at 200 s in the main probe and at 409 s in a bonus run with max_tokens=16384 and a 600-second timeout. Only 4B active parameters at inference; 12-digit chain-of-thought does not finish inside a usable budget.
  • A4 — Qwen 3.8 27B. Returns a wrong answer (off by 300 thousand on a 12-digit number) rather than timing out. With reasoning on and 8192 token budget, the model emits confidently and incorrectly.

Qwen 3.6 ON crosses every wall. Gemma-on crosses every wall except A4. Qwen 3.8 ON crosses every wall except A4.

What still breaks at smaller sizes (or without reasoning)

  • Chain depth (3+ steps). Granite 1/4, Phi-4 2/4, all three no-reasoning 27B/26B-MoE configs 2-3/4. Both reasoning-on dense configs 4/4.
  • Named-variable chains. Same shape: Granite 1/4, Phi-4 1/4, no-reasoning 2-3/4, reasoning-on 4/4.
  • Large mixed-sign sums. Granite, Phi-4, Gemma-off all miss. Both Qwen-on configs and Gemma-on 4/4.
  • Mixed-unit arithmetic. Granite 1/4, Phi-4 2/4, Qwen-off 3/4. Reasoning-on configs and Gemma-off 4/4.
  • 12-digit +. Granite, Phi-4, Gemma-off, and both Qwen 3.8 configs miss. Qwen 3.6 off and Qwen 3.6 on solve it. Gemma-on times out.

Below reasoning-on-Qwen 3.6, every one of these is a "verify or use a tool" case. Reasoning on a large dense model is what crosses them.

The MoE surprise

Gemma 4 26B-A4B has 26B total parameters but only 4B active at inference. With reasoning off, it scores 31/36 — beating both 27B dense Qwens (Qwen 3.6 28, Qwen 3.8 27) at sub-second latency. With reasoning on, it lands at 35/36 (97.2%) at about 8 seconds per task. The MoE reasoning latency is much lower than the dense 27B reasoning latency measured via LM Studio (~40 s for Qwen 3.6; Gemma-on at ~15 s per task). Qwen 3.8 via Ollama comes in at ~6 s per task. Most of that gap is LM Studio vs Ollama, not the model. Reasoning-on dense remains slow.

The MoE wins on reasoning-off latency and accuracy. The dense 27B Qwen 3.6 wins on the last hard items. A4 is the 4B-active ceiling.

Practical readout

If you are running local models and you need them to do math:

  • Decimals, percent, currency, word problems: any model. Don't burn reasoning budget here.
  • Two- to four-digit arithmetic: any model.
  • Mixed-unit arithmetic (cm + mm, kg + g): Gemma 4 26B-A4B or Qwen 3.6 / 3.8 27B with reasoning on. Phi-4 and Granite 4.1 8B fail at this.
  • Multi-step chains with named variables, 3+ steps: the three reasoning-on configurations handle all four cases. Below that, verify.
  • Large mixed-sign sums (7+ operands): the three reasoning-on configurations solved this. Below that, verify or use a tool.
  • 12-digit addition: Qwen 3.6 ON solves it exactly. Qwen 3.8 ON misses by 300 thousand. Gemma-on times out. Below that, verify.
  • Multiplication past 3 digits: all three reasoning-on configurations solved it exactly (5,031,793,793). Below that, verify.

If you are deciding between 8B / 14B / 27B and care about math:

  • The 8B Granite is fast and accurate on the easy half.
  • The 14B Phi-4 buys ~11 percentage points over the 8B on the full probe.
  • The 27B Qwen 3.6 with reasoning on solves the entire probe. Adds ~40 s of latency per task (LM Studio backend). The right choice when accuracy matters and latency is fine.
  • The 27B Qwen 3.8 with reasoning on solves 35/36 at ~6 s per task (Ollama backend). The latency gap to Qwen 3.6 is mostly the backend, not the model. At the cost of 12-digit accuracy.
  • The 26B MoE Gemma without reasoning lands between the 14B and the reasoning-on 27B on score, at sub-second latency. The math-strongest non-reasoning option in the probe.
  • The 26B MoE Gemma with reasoning on lands at 35/36 (97.2%) at ~15 s per task (LM Studio backend). The 27B Qwen 3.8 with reasoning on also lands at 35/36 (97.2%) at ~6 s per task (Ollama backend). Pick by backend: under LM Studio, Gemma-on is the fastest 35/36; under Ollama, Qwen 3.8-on is.

If your pipeline calls the model hundreds of times in a loop, the reasoning-on 27B is not the right choice. The reasoning-off 26B-A4B or the reasoning-off 27B Qwen is.

If your pipeline calls the model once per document and accuracy matters, the reasoning-on Qwen 3.6 is the right choice.

Probe details

  • 36 tasks, all original numbers
  • temperature=0, top_p=1
  • Reasoning-off and non-reasoning runs: max_tokens=4096
  • Reasoning-on runs: max_tokens=8192
  • Same system prompt: "you are a careful arithmetic calculator, give only the final number"
  • LM Studio OpenAI-compatible endpoint at http://localhost:1234/v1 for Qwen 3.6 and Gemma
  • Ollama OpenAI-compatible endpoint at http://localhost:11434/v1 for Granite 4.1 8B and Qwen 3.8
  • Reasoning models with default LM Studio settings burn their entire token budget on the internal scratchpad without emitting a final answer. For automated extraction pipelines, set reasoning_effort: "none" to disable the scratchpad, or raise max_tokens to 8192+ to give it room to complete.
  • Raw responses + grades: tmp/math-probe/results-*.json
  • Runner + analyzer: tmp/math-probe/run_probe.py, analyze.py

Bottom line

Size helps on the easy half of the math probe — 8B → 14B → 27B moves 20 → 24 → 28 with reasoning off. Reasoning on top of size crosses the rest: Qwen 3.6 27B with reasoning on jumps from 28 to 36. The MoE with only 4B active parameters comes within one task of the dense 27B at a quarter of the reasoning latency. The newer Qwen 3.8 is one task behind. Pick by backend and latency budget.