Local Model Bench
  • Home
  • Cases
  • Methodology
  • About

benchmarks

A collection of 1 post
Eight empty chairs around a candle-lit tavern table strewn with face-down cards; two flipped cards reveal wolf silhouettes
benchmarks

We made 8 local LLMs play Werwolf. The liars won.

We built a Werwolf game master and sat eight local LLMs around one village: 25 games, ~2,200 model calls. One model won 83% of its wolf games and never cast a bad ballot. One kept outing its own secret role. And the 135M underdog confessed to everything — including being a villager.
04 Oct 2026 5 min read
Page 1 of 1
Local Model Bench © 2026
  • Impressum
  • Privacy / Datenschutz
Powered by Ghost