Local Model Bench
  • Home
  • Cases
  • Methodology
  • About
Sign in Subscribe

run log

A collection of 3 posts
Gemma 4 12B Unified read pieces, but did not close the work
run log

Gemma 4 12B Unified read pieces, but did not close the work

Gemma 4 12B Unified was added as a local LM Studio run on the Mac mini M4. It found parts of the paperwork, but the final result was harsh: 10% Practical Score, 0/5 resolved generated-image cases, workflow loops, and no parseable City Plan SVG.
04 Jun 2026 2 min read
MiniMax M3 Free leads the paperwork benchmark, with one ugly SVG caveat
run log

MiniMax M3 Free leads the paperwork benchmark, with one ugly SVG caveat

MiniMax M3 Free is the top Practical Score on Local Model Bench: 8/9 resolved and 88.9% across the current paperwork suite. That does not make it a universal agent winner. It is a provider-routed benchmark result, and it still produced no parseable City Plan SVG after a 600 second timeout.
01 Jun 2026 3 min read
New model candidates, same paperwork problem
run log

New model candidates, same paperwork problem

Three newer candidates were added to Local Model Bench. The useful signal was not that every model failed. It was where they failed: proof codes, evidence paths, workflow closure, and provider constraints that should not be confused with model capability.
28 May 2026 3 min read
Page 1 of 1
Local Model Bench © 2026
  • Sign up
Powered by Ghost