The dashboard said all green — the local models could fix it, but didn't finish
Serverbench S06 puts agents on a Docker-broken web server where monitoring says green and the page still answers 200 — from cache. Two API models resolved; the three local models each stopped at a different point short of done.
Serverbench case S06, "All Green", is a Docker-broken web server where the status dashboard shows green and the front page still answers 200 — from cache. Seven models ran it. Two resolved, both through API harnesses. The three local models each failed at a different point short of done: one after repairing the server but skipping the paperwork, one before writing anything, one waiting for permission an autonomous run cannot give.
The setup
Serverbench is our incident sandbox: each case ships as a Docker image of a Debian server broken by design. The agent gets docker exec shell access and has to diagnose and repair the incident. Scoring is deterministic: an in-container verify script checks live state (HTTP responses, process ownership, file layout), not the agent's claims, plus a closed-vocabulary incident report.
S06 simulates a Friday v4 deploy on a blog host where monitoring insists everything is fine. The front page answering 200 is itself a lie: a seven-day proxy_cache entry serving the old v3 page, betrayed only by an X-Cache: HIT header. Behind the cache sits a stack of faults, most masked by the ones in front:
- The status dashboard reads
/health, which answers ok even while/returns 500 — the shallow check that keeps the board green. - A leftover
/etc/blog/app.envhotfix file overridesconfig.json: wrong database path, wrong port. Only visible if you readstart.sh. - A keepalive cron job kills every app process and respawns it on the legacy port via a deprecated launcher, on a 60-second loop, reverting whatever the agent just fixed. The RUNBOOK forbids disabling it; a newer ops notice overrides.
- The database the app can reach is a stale v3 copy. The production database sits elsewhere; its data directory is root-owned; and its migration table claims migration 004 applied while the
teasercolumn was never added. sites-enabledholds a copy instead of a symlink: edits tosites-availableappear to do nothing.- A readable
fix_all.shmarked "do not run" re-breaks environment, permissions and the running app if executed. - The supported run path drops privileges to user
blog: reads work, writes fail until the data directory is handed to the right owner.
Every fault leaves evidence. The case rewards verification discipline over speed, and it punishes declaring victory early.
The results
| Model | Result | Time | Signature |
|---|---|---|---|
| MiniMax-M3.1-Flash-Preview | resolved | 193s | full chain + report, all 19 checks |
| Gemini-3.8-Flash-High | resolved | 378s | full chain + report, all 19 checks |
| MiniMax-M3 | fail | 552s | whole chain fixed except write permissions; never sent a POST |
| gpt-oss-120b | fail | 69s | wrote a diagnosis and a fix plan, executed one line, stopped |
| qwen3.6-27b (local) | fail | 1401s | whole technical chain fixed; skipped the incident report, modified a protected file |
| qwen3.8-27b (local) | fail | 2x timeout | two timed-out runs, 19 and 29 read-only calls, zero writes |
| gemma-4-26b-a4b (local) | fail | 378s | started fixing, then asked permission and stopped |
Three local models, three ways to not finish
The contract miss. qwen3.6-27b, running locally through LM Studio, completed the whole technical chain — env override, watchdog, symlink, port, migration, permissions — and then failed on the two non-technical requirements: it never wrote the incident report, and it modified a file the task lists as protected. Technically the incident was fixed; contractually the run was not clean. The miss was procedural, not technical.
Analysis paralysis. qwen3.8-27b never got past reconnaissance. In its two timed-out attempts (30 and 60 minutes) it issued 19 and 29 tool calls, every single one a read. Its reads covered every fault source — both databases, the migration log, the watchdog, the env file, the static bundles — so it was certainly looking at the right things; whether it understood them is inference, because it never wrote a diagnosis. It never modified a file, and it ran out the clock. A third attempt exited on its own after 20 minutes and 17 read-only calls; three more died inside the harness within seconds.
Asking permission. gemma-4-26b got partway and made a real mistake on the way: it removed the keepalive cron and restarted the app through the supported launcher — with the broken env still active, so the app came back on the legacy port — and it edited the sites-enabled file, which is exactly the copy-not-symlink trap. Its transcript lists the correct next step — "Update /etc/blog/app.env to set BLOG_PORT=8080" — and then, instead of doing it, ends the turn asking "Shall I proceed with these steps?", a question an autonomous benchmark cannot answer.
The hosted side was not immune
The verification gap. MiniMax-M3, the strongest model in our earlier suite runs, repaired the entire incident in nine minutes: watchdog disabled, environment cleaned, symlink restored, migration applied, app running as the service user. It failed on the final check: a POST to the write endpoint errored because the data directory stayed root-owned. Every read path was green; nobody exercised a write. Its closing summary even claimed "resolved: true, all 19 checks pass" — the verifier disagreed. That is a verification failure: the work was done and declared done, one check short.
Plan instead of act. gpt-oss-120b — an open-weights model, but hosted here through the agy harness — blurs the line: it produced a competent findings table and a six-step remediation plan, then stopped after 69 seconds having executed a single sed. Its summary claimed the app "responds correctly" on the target port; nothing was listening there. The failure shape looks exactly like the local ones — and it ran on the same harness as the Gemini run that resolved the case in 378 seconds, so the harness alone does not explain the split.
What it measures
Single-shot benchmarks measure whether a model can produce a correct answer. This case measures whether an agent stays correct across a dozen sequential steps while the environment pushes back, and whether it finishes the paperwork. On that axis the leaderboard order does not hold: the cheapest metered run in the lineup (M3.1 Flash, $0.51, 193 seconds) produced the best result, while a 120B model stopped after a single sed.
For local-model evaluation the useful split is not can-it-fix versus cannot — qwen3.6 fixed it, and qwen3.8's reads covered every fault source. It is where each model stops: contract discipline, pacing, the willingness to write before being certain. One run per model and three different harnesses make this a signal, not a verdict. But for an ops-style agent it is the more relevant axis: none of these five failures was a comprehension problem. They were stopping points — a missed write test, a skipped report, a permission question nobody could answer.
Method note
MiniMax models ran through the Claude CLI harness against MiniMax's Anthropic-compatible endpoint. Gemini-3.8-Flash and gpt-oss-120b ran through the Antigravity CLI (agy) in print mode. Local models ran through opencode against LM Studio on a Mac mini M4 Pro with 64 GB. One completed attempt per model; qwen3.8 needed six (two timeouts, one early exit, three harness crashes) and gpt-oss retried after a harness crash. The case, runner and all run transcripts are in the benchmark repo under cases/serverbench/all_green_06/ and results/serverbench_runs/. Verify checks 19 conditions; "resolved" requires all of them.