Reasoning budget beats raw size on local math
A 27B dense model with reasoning on solves all 36 math tasks. Reasoning budget beats raw size on local math. Tested 8 model configs across 9 stress dimensions.
Field note — Local Model Bench — September 2026
Short answer
A 27B dense model with reasoning on solves all 36 math tasks in this probe. The 27B dense model one generation newer, with reasoning on, solves 35. A 26B MoE with reasoning on (only 4B active parameters) also solves 35 of 36 — only the 12-digit sum times out. Switch reasoning off on those same models and the dense 27B drops to 28 of 36, the 4B-active MoE drops to 31. Drop the size and an 8B solves 20, a 14B solves 24.
Size helps. Reasoning helps more. They interact, and they only get you to 100% when both are stacked on the right architecture — and the right configuration matters as much as the right model. Generation does not help uniformly: the newer Qwen loses ground on 12-digit addition that the older Qwen solved.
What the probe measures
36 tasks across 9 stress dimensions, four tasks per dimension. Original numbers throughout — no GSM8K, no MATH, no public benchmark contamination. Tolerance is exact or ±0.01 / ±0.1 / ±1.0 depending on the task.
| Dimension | What scales |
|---|---|
| Digit count | digits per operand (2 → 4 → 6 → 12) |
| Operand count | operands to sum (2 → 5 → 10 → 20) |
| Chain depth | sequential operations (1 → 2 → 3 → 5 steps) |
| Decimal precision | cents → 4-decimal sum |
| Signed numbers | clean → 1 negative → mixed signs → large mixed |
| Word problems | with irrelevant numbers in the prompt |
| Unit conversion | as part of the arithmetic |
| Named variables | dependency chain (1 → 2 → 3 → 5 steps) |
| Multiplication size | 1×1 → 5×5 digits |
Models and configurations
| Model | Size | Architecture | Reasoning | Score |
|---|---|---|---|---|
| Granite 4.1 8B | 8B | dense | no | 20/36 (55.6%) |
| Phi-4 | 14B | dense | no | 24/36 (66.7%) |
| Qwen 3.6 27B | 27B | dense | off | 28/36 (77.8%) |
| Qwen 3.8 27B | 27B | dense | off | 27/36 (75.0%) |
| Gemma 4 26B-A4B | 26B (4B active) | MoE | off | 31/36 (86.1%) |
| Gemma 4 26B-A4B | 26B (4B active) | MoE | on | 35/36 (97.2%) |
| Qwen 3.8 27B | 27B | dense | on | 35/36 (97.2%) |
| Qwen 3.6 27B | 27B | dense | on | 36/36 (100.0%) |
temperature=0, top_p=1. Reasoning-on runs used max_tokens=8192; the rest used max_tokens=4096. Same system prompt for all: "you are a careful arithmetic calculator, give only the final number."
Per-dimension accuracy
| Dimension | Granite 8B | Phi-4 14B | Qwen 3.6 off | Qwen 3.8 off | Gemma MoE off | Gemma MoE on | Qwen 3.8 on | Qwen 3.6 on |
|---|---|---|---|---|---|---|---|---|
| Digit count | 2/4 | 3/4 | 4/4 | 3/4 | 3/4 | 3/4 | 3/4 | 4/4 |
| Operand count | 3/4 | 3/4 | 3/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 |
| Chain depth | 1/4 | 2/4 | 2/4 | 2/4 | 3/4 | 4/4 | 4/4 | 4/4 |
| Decimal precision | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 |
| Signed numbers | 2/4 | 3/4 | 3/4 | 2/4 | 3/4 | 4/4 | 4/4 | 4/4 |
| Word problems | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 |
| Unit conversion | 1/4 | 2/4 | 3/4 | 3/4 | 4/4 | 4/4 | 4/4 | 4/4 |
| Named variables | 1/4 | 1/4 | 2/4 | 2/4 | 3/4 | 4/4 | 4/4 | 4/4 |
| Multiplication | 2/4 | 2/4 | 3/4 | 3/4 | 3/4 | 4/4 | 4/4 | 4/4 |
Every model is perfect on decimal precision and word problems. On multiplication, the two generation-different Qwens tie at 3/4 with reasoning off; all three reasoning-on configurations solve it cleanly.
The same model, two configurations (Qwen 3.6)
The cleanest signal in the probe is comparing Qwen 3.6 27B to itself:
- Reasoning off (max_tokens=4096): 28/36 (77.8%), 1.47 s per task.
- Reasoning on (max_tokens=8192): 36/36 (100.0%), ~38 s per task.
Same weights, same temperature, same system prompt. The only difference is whether the model uses its internal scratchpad channel before emitting a final answer. The 26× slowdown is real.
Generation drift (Qwen 3.6 vs Qwen 3.8)
The two 27B dense models differ on two tasks:
- A4 — 12-digit + 12-digit + 12-digit. Qwen 3.6 ON solves it exactly. Qwen 3.8 ON misses by 300 thousand (1585059803977 vs 1585059503977). Same in the reasoning-off configuration: 3.6 solves it exactly, 3.8 misses by 40 thousand (1585059464077).
- I4 — 84736 × 59382. Qwen 3.8 ON solves it exactly. Qwen 3.6 ON also solves it exactly. In the reasoning-off configuration, Qwen 3.6 is off by 62 thousand and Qwen 3.8 is off by 92 thousand — both within 0.002% of the true product, both close, and 3.6 marginally closer.
What still breaks with reasoning on
After reasoning is on, two models have exactly one wall left each:
- A4 — Gemma 4 26B-A4B. Times out at 200 s in the main probe and at 409 s in a bonus run with
max_tokens=16384and a 600-second timeout. Only 4B active parameters at inference; 12-digit chain-of-thought does not finish inside a usable budget. - A4 — Qwen 3.8 27B. Returns a wrong answer (off by 300 thousand on a 12-digit number) rather than timing out. With reasoning on and 8192 token budget, the model emits confidently and incorrectly.
Qwen 3.6 ON crosses every wall. Gemma-on crosses every wall except A4. Qwen 3.8 ON crosses every wall except A4.
What still breaks at smaller sizes (or without reasoning)
- Chain depth (3+ steps). Granite 1/4, Phi-4 2/4, all three no-reasoning 27B/26B-MoE configs 2-3/4. Both reasoning-on dense configs 4/4.
- Named-variable chains. Same shape: Granite 1/4, Phi-4 1/4, no-reasoning 2-3/4, reasoning-on 4/4.
- Large mixed-sign sums. Granite, Phi-4, Gemma-off all miss. Both Qwen-on configs and Gemma-on 4/4.
- Mixed-unit arithmetic. Granite 1/4, Phi-4 2/4, Qwen-off 3/4. Reasoning-on configs and Gemma-off 4/4.
- 12-digit +. Granite, Phi-4, Gemma-off, and both Qwen 3.8 configs miss. Qwen 3.6 off and Qwen 3.6 on solve it. Gemma-on times out.
Below reasoning-on-Qwen 3.6, every one of these is a "verify or use a tool" case. Reasoning on a large dense model is what crosses them.
The MoE surprise
Gemma 4 26B-A4B has 26B total parameters but only 4B active at inference. With reasoning off, it scores 31/36 — beating both 27B dense Qwens (Qwen 3.6 28, Qwen 3.8 27) at sub-second latency. With reasoning on, it lands at 35/36 (97.2%) at about 8 seconds per task. The MoE reasoning latency is much lower than the dense 27B reasoning latency measured via LM Studio (~40 s for Qwen 3.6; Gemma-on at ~15 s per task). Qwen 3.8 via Ollama comes in at ~6 s per task. Most of that gap is LM Studio vs Ollama, not the model. Reasoning-on dense remains slow.
The MoE wins on reasoning-off latency and accuracy. The dense 27B Qwen 3.6 wins on the last hard items. A4 is the 4B-active ceiling.
Practical readout
If you are running local models and you need them to do math:
- Decimals, percent, currency, word problems: any model. Don't burn reasoning budget here.
- Two- to four-digit arithmetic: any model.
- Mixed-unit arithmetic (cm + mm, kg + g): Gemma 4 26B-A4B or Qwen 3.6 / 3.8 27B with reasoning on. Phi-4 and Granite 4.1 8B fail at this.
- Multi-step chains with named variables, 3+ steps: the three reasoning-on configurations handle all four cases. Below that, verify.
- Large mixed-sign sums (7+ operands): the three reasoning-on configurations solved this. Below that, verify or use a tool.
- 12-digit addition: Qwen 3.6 ON solves it exactly. Qwen 3.8 ON misses by 300 thousand. Gemma-on times out. Below that, verify.
- Multiplication past 3 digits: all three reasoning-on configurations solved it exactly (5,031,793,793). Below that, verify.
If you are deciding between 8B / 14B / 27B and care about math:
- The 8B Granite is fast and accurate on the easy half.
- The 14B Phi-4 buys ~11 percentage points over the 8B on the full probe.
- The 27B Qwen 3.6 with reasoning on solves the entire probe. Adds ~40 s of latency per task (LM Studio backend). The right choice when accuracy matters and latency is fine.
- The 27B Qwen 3.8 with reasoning on solves 35/36 at ~6 s per task (Ollama backend). The latency gap to Qwen 3.6 is mostly the backend, not the model. At the cost of 12-digit accuracy.
- The 26B MoE Gemma without reasoning lands between the 14B and the reasoning-on 27B on score, at sub-second latency. The math-strongest non-reasoning option in the probe.
- The 26B MoE Gemma with reasoning on lands at 35/36 (97.2%) at ~15 s per task (LM Studio backend). The 27B Qwen 3.8 with reasoning on also lands at 35/36 (97.2%) at ~6 s per task (Ollama backend). Pick by backend: under LM Studio, Gemma-on is the fastest 35/36; under Ollama, Qwen 3.8-on is.
If your pipeline calls the model hundreds of times in a loop, the reasoning-on 27B is not the right choice. The reasoning-off 26B-A4B or the reasoning-off 27B Qwen is.
If your pipeline calls the model once per document and accuracy matters, the reasoning-on Qwen 3.6 is the right choice.
Probe details
- 36 tasks, all original numbers
temperature=0,top_p=1- Reasoning-off and non-reasoning runs:
max_tokens=4096 - Reasoning-on runs:
max_tokens=8192 - Same system prompt: "you are a careful arithmetic calculator, give only the final number"
- LM Studio OpenAI-compatible endpoint at
http://localhost:1234/v1for Qwen 3.6 and Gemma - Ollama OpenAI-compatible endpoint at
http://localhost:11434/v1for Granite 4.1 8B and Qwen 3.8 - Reasoning models with default LM Studio settings burn their entire token budget on the internal scratchpad without emitting a final answer. For automated extraction pipelines, set
reasoning_effort: "none"to disable the scratchpad, or raisemax_tokensto 8192+ to give it room to complete. - Raw responses + grades:
tmp/math-probe/results-*.json - Runner + analyzer:
tmp/math-probe/run_probe.py,analyze.py
Bottom line
Size helps on the easy half of the math probe — 8B → 14B → 27B moves 20 → 24 → 28 with reasoning off. Reasoning on top of size crosses the rest: Qwen 3.6 27B with reasoning on jumps from 28 to 36. The MoE with only 4B active parameters comes within one task of the dense 27B at a quarter of the reasoning latency. The newer Qwen 3.8 is one task behind. Pick by backend and latency budget.