Gemini 3.8 Flash through Antigravity: 9 of 9 hidden-oracle passes, but visible_checks fail on every image case
27/27 hidden-oracle passes — the headline. Visible_checks fail on every P-case at every effort: strict 4/9 (was 9/9), practical 72.2% (was 100%), rank #6 tied with qwen3.6-27b on score.
Gemini 3.8 Flash through Antigravity: 9 of 9 hidden-oracle passes, but visible_checks fail on every image case
Gemini 3.8 Flash shipped on September 2, 2026. Google's launch framing was specific: "engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows." That is the kind of claim Local Model Bench can stress-test directly.
The current nine-case paperwork suite is not glamorous. It hands a model a synthetic private-document folder — invoice scans, vendor tables, bank exports, purchase orders, protected source files — and asks it to produce a finished audit artifact. The verifier runs four common checks: audit_result_exists, visible_checks_pass, core_oracle_pass, and hidden_oracle_pass. A case is "strict resolved" only if all four return true; it is "core" if core_oracle_pass is true. Practical Score is 0.5 × strict + 0.5 × core.
So we ran Gemini 3.8 Flash through Google's own Antigravity CLI (agy) on the Mac mini M4, against all nine cases, at all three reasoning effort levels (high, medium, low). The headline is short and unambiguous: every one of the 27 runs passed the strict hidden oracle. The honest second finding is just as short: visible_checks_pass failed on all five P-cases (P01–P05) at every effort level. The strict score is therefore 4/9 (the four W-cases fully resolve), core is 9/9 (core_oracle passes everywhere), and the practical lands at 72.2%. Reasoning effort still moved the wall-clock time but not the score.
What ran
Gemini 3.8 Flash is available through the Google Antigravity CLI (agy). The CLI exposes three reasoning efforts: high, medium, and low.
We ran each of the nine Local Model Bench paperwork cases at all three effort levels, for 27 runs total:
- Five generated-image cases (P01–P05): scanned invoices plus CSV cross-checks
- Four agentic workflow cases (W04–W07): messy intake folder, email-attachment revisions, remittance split, credit offset
- Same
model_prompt.md, sameincoming/directory, same verifier, same hidden oracle
The model had to read the scans, choose the active sources, ignore the stale files, write the artifacts into work/, leave the incoming/ folder byte-identical, and emit a proof.txt that matches a hidden numeric code.
The headline table
| Effort | Practical Score | Strict Resolved | Hidden Oracle | Core | Visible Checks |
|---|---|---|---|---|---|
| gemini-3.8-flash-high | 72.2% | 4/9 | 9/9 | 9/9 | 4/9 |
| gemini-3.8-flash-medium | 72.2% | 4/9 | 9/9 | 9/9 | 4/9 |
| gemini-3.8-flash-low | 72.2% | 4/9 | 9/9 | 9/9 | 4/9 |
| Reference: opencode/minimax-m3-free (rank #1) | 88.9% | 8/9 | 8/9 | 8/9 | 8/9 |
| Reference: qwen3.6-27b (best local) | 72.2% | 5/9 | 8/9 | 8/9 | 5/9 |
Same nine cases. Same verifier. Same four checks. The strict-resolved and hidden-oracle columns diverge on antigravity because visible_checks_pass fails on all five image cases — that is the visible-format gate, not the model knowing the right answer. The hidden-oracle proof_code still matches on all nine.
Where the agent loop helped — and where it did not
The workflow cases (W04–W07) are the cleanest test of the agentic-first claim. They require:
- Reading the
incoming/folder before doing anything - Writing
work/normalized_manifest.json,work/document_index.json, three normalized text files,audit_result.json, andproof.txt - Keeping
incoming/byte-identical (the verifier checks this) - Producing a
proof_codethat matches the hidden oracle
On every workflow case at every effort level, Antigravity produced all artifacts on the first pass with audit_result_exists, visible_checks_pass, core_oracle_pass, and hidden_oracle_pass all returning true. That is the strictest possible score on a workflow case — full strict resolve — and it held across the 12 workflow runs.
The image cases (P01–P05) are a different shape, and this is where the format-sensitive failure lives. The original run_paperwork_v3_case_once.py runner calls the model once with a multimodal prompt — base64-encoded scan PNGs plus the CSV cross-checks baked into the user message — and expects a single parseable JSON object back. No tool calls. No file writes. No agent loop.
We did not have a CLI-native way to feed base64 images to the Antigravity CLI, so we gave the agent the case workspace directly: a folder with scans/*.png plus bank_export.csv, vendor_master.csv, purchase_orders.csv, and the task notes. We asked it to write audit_result.json. The Antigravity CLI read the scans, parsed the CSVs, and produced an audit artifact whose hidden-oracle proof_code matched the expected value on every P-case at every effort. That is a real, useful answer.
What it did not do is produce the exact visible_checks_pass shape the original single-shot multimodal runner expects. The hidden oracle — which validates the actual content, normalized files, and proof code — passes. The visible_checks runner, which validates the exact shape of the single-shot JSON response, does not.
This is not a knowledge failure. It is a shape mismatch between an agentic loop producing files on disk and a single-shot runner expecting an in-line JSON object. The two harnesses want different surfaces from the same model. Antigravity's strength — autonomous multi-step work — is exactly what the original multimodal runner was not designed to validate.
Where the effort level mattered
This is the most interesting finding, and it inverts the usual expectation.
Reasoning effort is the most expensive knob on Gemini 3.8 Flash. Antigravity's --effort flag controls it directly. The natural assumption is that higher effort means more careful answers, which means more cases resolved, which means a higher score.
That is not what the data shows.
The aggregate wall-clock time per effort, summed across the nine cases:
| Effort | Workflow cases (W04–W07) | Image cases (P01–P05) | Total |
|---|---|---|---|
| High | 92 + 186 + 207 + 180 = 665s | first-run setup + 30 + 76 + 45 + 45 ≈ ~525s | ~1,190s (~20 min) |
| Medium | 90 + 150 + 135 + 120 = 495s | 45 + 45 + 30 + 45 + 45 = 210s | ~705s (~12 min) |
| Low | 75 + 136 + 75 + 60 = 346s | 30 + 30 + 45 + 30 + 45 = 180s | ~526s (~9 min) |
(One P01 high run was killed by the watchdog at 5:30 because the first call to a new model through Antigravity carries cold-start overhead. The audit artifact itself was complete and matched the oracle; the run only stopped because the CLI did not return cleanly after writing the file. The other four P-cases at high effort finished in 30–76 seconds.)
Across the matrix:
- The score is identical at every effort level. 4/9 strict, 9/9 core, 72.2% practical.
- The wall-clock time is ~2.3x faster at low than at high.
- The image cases barely change across effort. The workflow cases do — workflow runs at high are roughly 1.9x slower than at low, because the agent reads more files before settling on the action plan.
In other words, on this specific nine-case paperwork suite, the extra reasoning at high is paying for time, not for correctness.
This is the opposite of how the marketing framing talks about Gemini 3.8 Flash. Google's positioning treats high-effort reasoning as the default expectation for agentic work. The bench says the extra reasoning has no purchase on these nine cases. On this suite the hidden-oracle and core scores are already at the ceiling; extra effort has nothing left to buy in strict terms either.
What this means for the leaderboard
Two ways to read this.
The narrow read is that the leaderboard does not have a new #1. Antigravity lands at rank #6 with 72.2% practical — tied with qwen3.6-27b on score, but with a different profile: antigravity has 4/9 strict and 9/9 core, qwen3.6-27b has 5/9 strict and 8/9 core. The current top is still opencode/minimax-m3-free at 88.9% (8/9 strict, 8/9 core). What Antigravity adds to the leaderboard is the first reference row where the hidden oracle is universally clean (9/9) and the visible-check format gate is uniformly failing (4/9 visible) — a profile that does not exist anywhere else in the matrix today.
The wider read is what the leaderboard is actually measuring at this point. The cases in this suite were chosen to separate "smart" from "work-ready." When a model-plus-harness combination matches the hidden oracle on every case but fails the visible-check shape on every image case, the cases are no longer separating model tiers — they are separating which surface the harness produces. The strict score has become a measurement of harness compatibility as much as model capability.
That is still useful. It just is not a benchmark of the model in isolation. It is a benchmark of the model plus harness. The next set of cases for this benchmark needs to be harder, and probably needs to be harness-aware on both sides — test the agentic surface, then test the single-shot surface separately, then compare.
This is not a model ranking
A few caveats before this gets overread.
First, this is a nine-case synthetic benchmark. It tests paperwork-shaped problems well. It does not test long-form writing, math, general chat, code reasoning, or any of the dozens of other things a modern model is asked to do. With 9/9 hidden oracle and 9/9 core on this suite, the strict score is the bottleneck, and the bottleneck is harness-compatibility — not model knowledge.
Second, Antigravity CLI is the local first-party agent surface for Gemini 3.8 Flash. Routing the same model through the raw Gemini API, through OpenRouter, or through a different agent harness would produce different traces, different artifacts, and different scores. The benchmark measures the combination, not the model in isolation. The hidden-oracle passes here are real; they would not necessarily transfer to a single-shot multimodal harness, which is the surface the original P-case runner was designed for.
Third, the practical score is half strict resolved and half core-oracle pass. A model that understands most of a case and misses one closure detail still gets partial credit. That is intentional. It separates "got the main idea" from "finished the job." Antigravity's 72.2% reflects exactly that split: it understood all nine (9/9 core, 9/9 hidden) but finished strictly only on the four workflow cases where the harness surface matched the verifier's expectations.
The benchmarks we did not run
For context, here are the public benchmarks Google's launch post and the Google Cloud developer guide cite for Gemini 3.8 Flash. We are not running these ourselves; we are noting which numbers Google publishes, with which harness, so you know where the agentic-paperwork claim sits relative to Google's own framing. Treat each row as a Google-side claim with the harness Google chose — independent reproductions may differ.
| Benchmark | Google's number | Harness / set | Notes |
|---|---|---|---|
| Terminal-Bench 2.1 | 90.8% (developer guide) vs 89.4% (launch table) vs 87.6% (Artificial Analysis) | Google Cloud run vs Google launch run vs AA independent | Three different numbers from three different harnesses. Use the Google Cloud developer-guide number if you are citing the Antigravity-style agentic configuration. |
| DeepSWE v1.1 | 73.7% | Logan Kilpatrick / launch table | Within ~0.3 points of Claude Opus 5 (74.0%). Vendor-claim; treat as "in the same band," not "board SOTA." |
| Vals Finance Agent v2 | 61.44% ± 0.13 | Independent benchmark | Best published number for the model. The paperwork suite here is similar in shape but not the same workload. |
| τ³-bench Banking | 38.1% | Google Cloud developer guide, Google run | Up from 30.9% on 3.7 Flash. Independent τ³ reproductions of 3.8 Flash often land around 45% — cite the guide number only as a Google-side measurement. |
| HLE (unverified) | 45.4% | Google Cloud developer guide | Down from 45.7% on 3.7 Flash. Note that HLE-Verified (54.9% for the model) is a different set; do not conflate. |
We include this table because the article references some of these numbers in the body. The numbers themselves are not under our control; the harness choices are.
What this means in practice
For local private-document work, the question has not been "is the model smart." Gemini 3.8 Flash is clearly smart — the hidden-oracle and core-oracle passes on all 27 runs confirm that. The question has been "does the agent loop close the folder without leaking or stalling, and does the harness surface match what the verifier expects."
On the four workflow cases (W04–W07), Antigravity is the first closed-API combination we have tested that passed strict resolve on the first pass at every effort, with no visible repair loop or token-limit stall in those 12 runs. The cost difference between high and low on this workload is roughly a 2x latency multiplier for identical output.
On the five image cases (P01–P05), Antigravity passes hidden and core at every effort — meaning the model knows the right answer and produces an audit artifact whose proof_code matches the oracle. It does not produce the single-shot JSON response shape that the original multimodal runner validates with visible_checks_pass. That is a harness-compatibility finding, not a model-quality finding.
The next question for this benchmark is harder cases. With 9/9 hidden oracle and 9/9 core across the entire matrix, the suite is at saturation on what the model actually knows. The next step is cases that test harness-shape assumptions more directly, or that test what happens when the agent loops or stalls, or that require multi-step reasoning with hidden state, or that exercise something the current nine do not.