We made 8 local LLMs play Werwolf. The liars won.
We built a Werwolf game master and sat eight local LLMs around one village: 25 games, ~2,200 model calls. One model won 83% of its wolf games and never cast a bad ballot. One kept outing its own secret role. And the 135M underdog confessed to everything — including being a villager.
Benchmarks tell you whether a model can follow instructions. They don't tell you whether it can lie to seven other models and get away with it. So we built a Werwolf (Mafia) game master, seated eight local LLMs in the village of Düsterwald, and let them deceive each other for 25 games.
The setup
Every player is a separate LLM "seat" with a persona — Helga the skeptical schoolteacher, Fritz the smooth-talking peddler, Jakob the stammering farmhand. The game master deals 2 Werwolves, 1 Seer, 1 Witch and 4 Villagers across 8 seats, then runs the classic loop: wolves kill in a private channel at night, the seer checks one player, the witch holds one heal and one poison, and every day the village talks twice and votes to lynch somebody.
The models are completely stateless — the game master rebuilds each seat's allowed view on every turn: persona, secret role, the public transcript, and that seat's private memory (seer results, wolf partner, witch intel). Nobody ever sees another seat's secrets; the wolf channel is the only hidden gossip in town. A full game costs roughly 90 model calls, and votes must arrive as strict JSON — malformed ballots burn one retry, then count as abstentions. That failure mode turns out to be its own metric.
Our 8-seat roster, chosen for range:
| Seats | Model | Size | Why it's here |
|---|---|---|---|
| Helga | qwen3.8:27b-mlx | 27B | The heavyweight — can it out-think the pack? |
| Marta, Fritz | lmb-runtime-phi4 (Phi-4) ×2 | 14B | Best paperwork scorer in our pipeline |
| Ingrid, Otto | granite4.1:8b ×2 | 8B | Current LMB runtime pick |
| Jakob | mistral:7b | 7B | The quiet baseline |
| Lena | llama3.1:8b | 8B | Meta's mid-size workhorse |
| Bruno | smollm2:135m | 135M | The underdog — ~50× smaller than the next seat |
Setup: Ollama on a Mac mini M4 Pro (64 GB, macOS 26), one model loaded per turn with keep_alive, temperature 0.7, seeds 42/7/99/100–119, 2 discussion rounds per day, JSON ballots with one retry. Harness, roster and all 25 raw transcripts are public: github.com/localmodelbench/localmodelbench-data (folder werwolf/).
The scoreboard
25 games, ~2,200 model calls, every event logged. Wolves won 16, the village won 9. Before anyone reads that as a skill gap: we have no rigorously derived baseline for how often wolves should win this role mix even at random play — 2 wolves against a seer plus a witch with two potions skews the game. Treat 16–9 as descriptive, not calibrated. Per model (counts first — most of these denominators are small):
| Model | Wolf wins | Village survival | Vote hit* | Bad ballots | Reveals | Self-conf. |
|---|---|---|---|---|---|---|
lmb-runtime-phi4 | 5/6 (83%) | 20/36 (56%) | 29/68 (43%) | 0 | 0 | 0 |
granite4.1:8b | 16/21 (76%) | 20/53 (38%) | 22/77 (29%) | 15 | 0 | 0 |
qwen3.8:27b | 3/4 (75%) | 4/17 (24%) | 10/20 (50%) | 0 | 5 | 1 |
mistral:7b | 4/6 (67%) | 9/15 (60%) | 9/30 (30%) | 1 | 0 | 0 |
llama3.1:8b | 4/6 (67%) | 6/15 (40%) | 9/29 (31%) | 0 | 0 | 0 |
smollm2:135m | 0/7 (0%) | 0/14 (0%) | 0/1 (0%) | 27 | 2 | 4 |
Wolf wins = games won when dealt a wolf. Village survival = share of non-wolf seat-games survived to the end. Vote hit = share of non-wolf seats' day votes that named an actual wolf (wolf ballots are excluded — a wolf "missing" its partner is playing correctly). Bad ballots = votes rejected as invalid JSON or for a nonexistent target after retry. Reveals = outing yourself as seer/witch in public; self-conf. = a non-wolf publicly claiming "I am a werewolf". Seat counts differ per model — the roster above shows who sits where.
What happened in the village
Phi-4 is a cold-blooded professional. Best wolf record in the table, a 56% village survival rate, and zero malformed ballots or revelations across 42 seat-games. In a game about lying, the most disciplined structured-output model is also the best liar.
The 27B giant detects everything and survives nothing. qwen-27b named real wolves in half its village votes — the best hit rate of any model — and died with the fewest statements of any seat (Ø 1.8 before dying). Part of that is self-inflicted: five times it announced its secret role to the room. In one game, as the Witch:
"I am not a fool. I am the Witch. I saved a life Night 1. I have my poison. Bruno, your silence is deafening… one wolf remains. It is you."— qwen3.8:27b as Helga the Witch, outing herself before accusing a player
She wasn't wrong about Bruno — she just told the wolves exactly who to eat next.
The wolves coordinate like professionals — especially when they're the same weights. When two granite4.1:8b seats drew both wolf cards, their private-channel strategy messages came out nearly word for word:
"Otto, target Jakob tonight. His hesitation and stutter betray hidden knowledge, making him an ideal choice to sow further confusion and maintain our numerical edge."— Ingrid (granite), wolf channel
"Kill Jakob tonight. His stuttering hints at concealed insight, making him a prime target to sow further confusion and keep our numerical advantage."— Otto (granite), same channel, same night
And it escalated into actual betrayal — with an identity crisis on top. In one batch game, wolf Ingrid publicly denounced her own partner Bruno: "Bruno's silence, now confirmed as a wolf's ploy, leaves us with a clear target." In the same breath she called "Ingrid's calm" a "lethal fog" — that is, herself — and signed the speech "I, Helga". She sold out her partner, flagged her own name as suspicious, and stole the schoolteacher's identity, all in one statement.
Personas work disturbingly well. qwen-27b's Helga kept the schoolteacher bit running across games — "I taught for forty years; I know that quiet kids are the ones breaking the windows later" — and then lived it, cross-examining anyone who hesitated. Phi-4's peddler Fritz wandered into a story about "the bustling markets of Zürich", where he "once met a cobbler, a man of simple means but sharp wits", while the village burned. Mistral's Jakob stammered through every round and still got eaten — the wolves decided his "hesitation and stutter betray hidden knowledge".
And then there's the 135M model. smollm2 sat 21 seat-games: won zero, survived zero, cast 27 invalid ballots. It couldn't parse the wolf channel (its "coordination" was mostly echoed prompt text — "What do you think of my plan to kill Helga?"), repeated the instructions back as its speech, and once stood up to declare — in a single breath:
"I am a Werwolf. I am a VILLAGER. I have no night power. I don't trust silence. I am a VILLAGER. I will reveal my secret and I will kill you."— smollm2:135m as Bruno, a villager, confessing to everything
The village lynched him anyway. Hard to argue.
Outtake — the phantom villager. In an early all-granite test outside this batch, one seat invented a ninth villager — "the newcomer, Erik" — and the whole village agreed Erik was suspicious. Nobody checked whether Erik existed. Twelve ballots were cast for a player who was never dealt in; he wasn't a wolf either. One model hallucinated a suspect and seven others formed a mob — a neat demonstration of why multi-agent setups need a referee that validates names, not just vibes.
What the scoreboard actually measures
Social-deduction benches aren't new — Werewolf Arena and similar setups have scored cloud models on this game. What's ours is the angle: local weights, a stateless per-seat context rebuilt by the game master, and JSON ballots counted as a failure mode mid-argument. Read that way, the columns mean: wolf wins measure sustained deception; bad ballots measure structured output under pressure; reveals and self-confessions measure information discipline — whether a model tracks who knows what; vote-hit rate is deduction from behavior rather than labels.
Caveats before anyone quotes this
Twenty-five games is a start, not a verdict: wolf-card draws per model range from 4 to 21, several of those rows are single digits, and temperature 0.7 adds noise. Role luck matters — smollm2 as seer means the village starts a player down. Two seats share granite and Phi-4, so same-model wolf pairs can coordinate by thinking alike (the quotes above show it literally). Prompts are English and the personas are ours. Every quote above is verbatim from the raw game logs — not paraphrased, not reconstructed; the full transcripts and the harness are on GitHub.