About this site

What's here

Three layers of content, distinguished by a coloured pill on each card:

  • Editorial (green) — opinionated thinking about the benchmark landscape. Not tied to a specific run.
  • Field notes (amber) — short observations from a single model run, no methodology update.
  • Run logs (blue) — longer accounts of one model, one case, the artifact that came out, and what broke.

Reference lines (Codex CLI, cloud APIs) appear without colour. They are not the same kind of artifact as a local run.

What's not here

  • No leaderboard games, no paid rankings, no monthly "model of the month".
  • No "10 best local LLMs" SEO content.
  • No sponsor brands on the home page.
  • No models tested that the author has not personally run on this hardware.

How it works

Every model listed here was actually run on the same hardware (a single Mac mini M4 Pro with 64GB unified memory, used for LM Studio local runs). Cloud API references (Codex, Gemini, GPT) are run on their respective APIs and priced in at zero local cost.

Each case has a strict pass/fail check on the final artifact — valid JSON, proof code present, SVG parses, evidence path matches. "Looks right" is not enough.

Full scoring rules, benchmark structure, and what's in and out of scope are on the Methodology page.

Source of truth

Case definitions and run logs live in a separate Astro repo on the same host. Numbers in articles match that repo at the time of writing. If a re-run on the same inputs gives different numbers, the article gets an addendum — it does not get quietly edited.

Me

One person, hobby budget, no team, no reviewers. Mistakes are listed in the change notes, not hidden. If something is wrong, the article says so. Subscribe if you want the next run log in your inbox — most of the work is published free.