Local Model Bench — Full Leaderboard

All 40 models that completed the current v1 paperwork suite, run thinking-enabled by default. Practical Score = 50% resolved cases + 50% core passes. Local LM Studio runs were executed on a Mac mini M4 with 64 GB unified memory.

← back to localmodelbench.com
OK near miss / core pass fail 9 cases · The Paperwork Trial
Type
Mode

40 of 40 shown

Rank Model Type Mode Practical Resolved Core Tried Case Matrix
1
kimi-k3-max via Devin CLI
reference provider 94.4% 8/9 9/9 9/9
2
swe-2-max via Devin CLI
reference provider 88.9% 8/9 8/9 9/9
3
minimax-m3-free via opencode
api cheap provider 88.9% 8/9 8/9 9/9
4
OpenAI GPT-5.5 (Codex CLI) codex-default
reference provider 83.3% 7/9 8/9 9/9
5
gpt-6-luna-high via Devin CLI
reference provider 77.8% 7/9 7/9 9/9
6
claude-opus-5-5-max via Devin CLI
reference provider 77.8% 7/9 7/9 9/9
7
OpenAI GPT-5.4 Mini (Codex CLI) gpt-5.4-mini
reference provider 77.8% 7/9 7/9 9/9
8
Muse-Glimmer† meta_muse-glimmer
local thinking 77.8% 6/9 8/9 9/9
9
grok-4-7-xhigh via Devin CLI
reference provider 72.2% 6/9 7/9 9/9
10
gpt-6-astra-max via Devin CLI
reference provider 72.2% 6/9 7/9 9/9
11
Ternary Bonsai 2 27B Thinking prism-ml_ternary-bonsai-2-27b-thinking
local thinking 72.2% 6/9 7/9 9/9
12
qwen3.6-27b† qwen_qwen3.6-27b
local thinking 72.2% 5/9 8/9 9/9
13
gemini-3.8-flash via Antigravity
reference provider 72.2% 4/9 9/9 9/9
14
gpt-6-luna-max via Devin CLI
reference provider 61.1% 5/9 6/9 9/9
15
gemma-4-26b-a4b† google_gemma-4-26b-a4b
local thinking 61.1% 4/9 7/9 9/9
16
qwen3.6-27b-thinking qwen_qwen3.6-27b-thinking
local thinking 61.1% 3/9 8/9 9/9
17
qwen3.8-27b openrouter_qwen_qwen3.8-27b
local no-think 55.6% 3/9 7/9 9/9
18
qwen3.6-35b-a3b† qwen_qwen3.6-35b-a3b
local thinking 38.9% 1/9 6/9 9/9
19
gemma-4-12b-thinking google_gemma-4-12b-thinking
local thinking 33.3% 1/9 5/9 9/9
20
Bonsai 27B† prism-ml_bonsai-27b
local thinking 33.3% 2/9 4/9 9/9
21
qwen3.6-flash openrouter_qwen_qwen3.6-flash
api cheap provider 33.3% 0/9 6/9 9/9
22
gemma-4-e4b† google_gemma-4-e4b
local thinking 27.8% 2/9 3/9 9/9
23
gemma-4-31b-it gemma-4-31b-it
local no-think 27.8% 0/9 5/9 9/9
24
gemini-3.1-flash-lite openrouter_google_gemini-3.1-flash-lite
api cheap provider 27.8% 0/9 5/9 9/9
25
gemini-2.5-flash openrouter_google_gemini-2.5-flash
api cheap provider 27.8% 0/9 5/9 9/9
26
Qwen3 VL 30B A3B openrouter_qwen_qwen3-vl-30b-a3b-instruct
api cheap provider 27.8% 0/9 5/9 9/9
27
Ternary Bonsai 2 27B prism-ml_ternary-bonsai-2-27b
local no-think 27.8% 0/9 5/9 9/9
28
qwen3-vl-32b-instruct openrouter_qwen_qwen3-vl-32b-instruct
api cheap provider 22.2% 0/9 4/9 9/9
29
Seed 2.0 Mini openrouter_bytedance-seed_seed-2.0-mini
api cheap provider 22.2% 0/9 4/9 9/9
30
Mistral Small 4 openrouter_mistralai_mistral-small-2603
api cheap provider 22.2% 0/9 4/9 9/9
31
mistral-small-3.2 mistralai_mistral-small-3.2
local no-think 16.7% 0/9 3/9 9/9
32
ministral-3-14b ministral-3-14b
local no-think 16.7% 0/9 3/9 9/9
33
gemma-4-12b-qat-thinking google_gemma-4-12b-qat-thinking
local thinking 16.7% 0/9 3/9 9/9
34
gemma-4-26b-a4b-nothink google_gemma-4-26b-a4b-nothink
local no-think 11.1% 0/9 2/9 9/9
35
gemma-4-12b google_gemma-4-12b
local no-think 5.6% 0/9 1/9 9/9
36
gemma-4-e2b† google_gemma-4-e2b
local thinking 0.0% 0/9 0/9 9/9
37
qwen3-vl-8b-instruct qwen3-vl-8b-instruct
local no-think 0.0% 0/9 0/9 9/9
38
qwen3-14b qwen_qwen3-14b
local no-think 0.0% 0/9 0/9 9/9
39
nemotron-3-nano-omni-30b-a3b-reasoning:free openrouter_nvidia_nemotron-3-nano-omni-30b-a3b-reasoning_free
api cheap provider 0.0% 0/9 0/9 9/9
40
ministral-3-3b mistralai_ministral-3-3b
local no-think 0.0% 0/9 0/9 9/9

Benchmark convention: thinking-enabled is the default. thinking = reasoning on, including rows marked † that were recorded under the previous no-think convention but produced reasoning tokens anyway (hover for totals). no-think = verified zero-reasoning runs, kept as the secondary condition; these models are being re-measured under thinking. provider = hosted API rows under provider defaults. See the methodology note.