All 40 models that completed the current v1 paperwork suite, run thinking-enabled by default. Practical Score = 50% resolved cases + 50% core passes. Local LM Studio runs were executed on a Mac mini M4 with 64 GB unified memory.
40 of 40 shown
| Rank | Model | Type | Mode | Practical | Resolved | Core | Tried | Case Matrix |
|---|---|---|---|---|---|---|---|---|
| 1 |
kimi-k3-max
|
reference | provider | 94.4% | 8/9 | 9/9 | 9/9 | |
| 2 |
swe-2-max
|
reference | provider | 88.9% | 8/9 | 8/9 | 9/9 | |
| 3 |
minimax-m3-free
|
api cheap | provider | 88.9% | 8/9 | 8/9 | 9/9 | |
| 4 |
OpenAI GPT-5.5 (Codex CLI)
|
reference | provider | 83.3% | 7/9 | 8/9 | 9/9 | |
| 5 |
gpt-6-luna-high
|
reference | provider | 77.8% | 7/9 | 7/9 | 9/9 | |
| 6 |
claude-opus-5-5-max
|
reference | provider | 77.8% | 7/9 | 7/9 | 9/9 | |
| 7 |
OpenAI GPT-5.4 Mini (Codex CLI)
|
reference | provider | 77.8% | 7/9 | 7/9 | 9/9 | |
| 8 |
Muse-Glimmer†
|
local | thinking | 77.8% | 6/9 | 8/9 | 9/9 | |
| 9 |
grok-4-7-xhigh
|
reference | provider | 72.2% | 6/9 | 7/9 | 9/9 | |
| 10 |
gpt-6-astra-max
|
reference | provider | 72.2% | 6/9 | 7/9 | 9/9 | |
| 11 |
Ternary Bonsai 2 27B Thinking
|
local | thinking | 72.2% | 6/9 | 7/9 | 9/9 | |
| 12 |
qwen3.6-27b†
|
local | thinking | 72.2% | 5/9 | 8/9 | 9/9 | |
| 13 |
gemini-3.8-flash
|
reference | provider | 72.2% | 4/9 | 9/9 | 9/9 | |
| 14 |
gpt-6-luna-max
|
reference | provider | 61.1% | 5/9 | 6/9 | 9/9 | |
| 15 |
gemma-4-26b-a4b†
|
local | thinking | 61.1% | 4/9 | 7/9 | 9/9 | |
| 16 |
qwen3.6-27b-thinking
|
local | thinking | 61.1% | 3/9 | 8/9 | 9/9 | |
| 17 |
qwen3.8-27b
|
local | no-think | 55.6% | 3/9 | 7/9 | 9/9 | |
| 18 |
qwen3.6-35b-a3b†
|
local | thinking | 38.9% | 1/9 | 6/9 | 9/9 | |
| 19 |
gemma-4-12b-thinking
|
local | thinking | 33.3% | 1/9 | 5/9 | 9/9 | |
| 20 |
Bonsai 27B†
|
local | thinking | 33.3% | 2/9 | 4/9 | 9/9 | |
| 21 |
qwen3.6-flash
|
api cheap | provider | 33.3% | 0/9 | 6/9 | 9/9 | |
| 22 |
gemma-4-e4b†
|
local | thinking | 27.8% | 2/9 | 3/9 | 9/9 | |
| 23 |
gemma-4-31b-it
|
local | no-think | 27.8% | 0/9 | 5/9 | 9/9 | |
| 24 |
gemini-3.1-flash-lite
|
api cheap | provider | 27.8% | 0/9 | 5/9 | 9/9 | |
| 25 |
gemini-2.5-flash
|
api cheap | provider | 27.8% | 0/9 | 5/9 | 9/9 | |
| 26 |
Qwen3 VL 30B A3B
|
api cheap | provider | 27.8% | 0/9 | 5/9 | 9/9 | |
| 27 |
Ternary Bonsai 2 27B
|
local | no-think | 27.8% | 0/9 | 5/9 | 9/9 | |
| 28 |
qwen3-vl-32b-instruct
|
api cheap | provider | 22.2% | 0/9 | 4/9 | 9/9 | |
| 29 |
Seed 2.0 Mini
|
api cheap | provider | 22.2% | 0/9 | 4/9 | 9/9 | |
| 30 |
Mistral Small 4
|
api cheap | provider | 22.2% | 0/9 | 4/9 | 9/9 | |
| 31 |
mistral-small-3.2
|
local | no-think | 16.7% | 0/9 | 3/9 | 9/9 | |
| 32 |
ministral-3-14b
|
local | no-think | 16.7% | 0/9 | 3/9 | 9/9 | |
| 33 |
gemma-4-12b-qat-thinking
|
local | thinking | 16.7% | 0/9 | 3/9 | 9/9 | |
| 34 |
gemma-4-26b-a4b-nothink
|
local | no-think | 11.1% | 0/9 | 2/9 | 9/9 | |
| 35 |
gemma-4-12b
|
local | no-think | 5.6% | 0/9 | 1/9 | 9/9 | |
| 36 |
gemma-4-e2b†
|
local | thinking | 0.0% | 0/9 | 0/9 | 9/9 | |
| 37 |
qwen3-vl-8b-instruct
|
local | no-think | 0.0% | 0/9 | 0/9 | 9/9 | |
| 38 |
qwen3-14b
|
local | no-think | 0.0% | 0/9 | 0/9 | 9/9 | |
| 39 |
nemotron-3-nano-omni-30b-a3b-reasoning:free
|
api cheap | provider | 0.0% | 0/9 | 0/9 | 9/9 | |
| 40 |
ministral-3-3b
|
local | no-think | 0.0% | 0/9 | 0/9 | 9/9 |
Benchmark convention: thinking-enabled is the default. thinking = reasoning on, including rows marked † that were recorded under the previous no-think convention but produced reasoning tokens anyway (hover for totals). no-think = verified zero-reasoning runs, kept as the secondary condition; these models are being re-measured under thinking. provider = hosted API rows under provider defaults. See the methodology note.