Google released Gemini 3.8 Flash on September 2, 2026. The same day, CellCog’s Agent Flash and Team Flash tiers moved to it. Nothing changes on your side: any chat or AI employee on a Flash tier is already running the new model, at the same pricing, with the same thinking setting it ran yesterday.
The release itself, Google’s full benchmark table, the pricing, and the restricted Cyber variant are covered on our Gemini 3.8 Flash launch post. This is the operator’s half.
On this page · 6 sectionsOpen
- Google released Gemini 3.8 Flash on September 2, 2026. The same day, CellCog’s two Flash tiers, Agent Flash and Team Flash, moved from Gemini 3.7 Flash to it.
- The operating point did not change: the same medium thinking setting, the same 1M-token context window, the same 64K output limit. Only the model id moved.
- Google held the price: $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, exactly what 3.7 Flash cost. Your pricing does not change either.
- Google’s own table puts 3.8 Flash ahead of 3.7 Flash on every published row, led by DeepSWE v1.1 (73.7% vs 65.3%), OSWorld-2.0 (59.0% vs 50.6%), and Terminal-bench 2.1 (89.4% vs 85.8%). Vendor numbers, labeled as such.
- The honest caveat, in Google’s words: 3.8 Flash ‘works harder’ and may use more tokens on complex tasks. Same per-token price, possibly more tokens per job. We kept the medium setting for exactly that reason and are measuring Flash cost per turn over the next week.
- Gemini 3.7 Flash stays in our catalog as a one-line rollback. Core, Max, and Creative tiers are unchanged.
- Nothing to do on your side: any chat or AI employee on a Flash tier is already running Gemini 3.8 Flash.
§ 01What changed
Two of CellCog’s tiers moved. Five did not.
Agent Flash and Team Flash moved from Gemini 3.7 Flash to Gemini 3.8 Flash. Everything else about their operating point is untouched: the same medium thinking setting we adopted on August 13 with 3.7 Flash, the same 1M-token context window, the same 64K output limit. The routing table that picks the model per request now names 3.8 Flash on exactly those two rows; the same model also took over the large-document reading we do behind the scenes.
Agent Core, Agent Max, Team Core, and Team Max run on Claude Fable 5.1, where they moved yesterday. Agent Creative runs on Claude Opus 5. Neither changed today. If you want the full picture of how modes and tiers fit together, it is in Flash, Core, Max.
§ 02What Gemini 3.8 Flash is
Google’s newest Flash-class model, and by Google’s count its third Flash release in six weeks. It is a drop-in successor to 3.7 Flash on paper: the developer guide lists the same 1,048,576-token context window, the same 65,536-token maximum output, and the same low, medium, and high thinking levels, with medium as the default. Google held the price too: $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, then $1.50 and $7.50 from January 1, 2027, the same schedule 3.7 Flash already had.
What moved is the benchmark column, most of all on the agent-shaped rows. These are Google’s numbers from the announcement, and we treat them that way.
| Benchmark | 3.8 Flash | 3.7 Flash |
|---|---|---|
| Terminal-bench 2.1 (agentic terminal coding) | 89.4% | 85.8% |
| DeepSWE v1.1 (long-horizon software engineering) | 73.7% | 65.3% |
| Vals Finance Agent v2 | 61.4% | 59.0% |
| OSWorld-2.0 (agentic computer use) | 59.0% | 50.6% |
| HLE-Verified (expert reasoning) | 54.9% | 53.6% |
| Terminal-bench 4.0 (general agent capabilities) | 19.1% | 11.2% |
| Gray Swan prompt injection, attack success within 15 attempts (lower is better) | 5.5% | 9.2% |
The last row in the table matters most to us. A Flash tier reads a lot of untrusted text: web pages, inbound email, documents somebody else wrote. Google charts 3.8 Flash at a 5.5% attack success rate on Gray Swan’s indirect prompt-injection benchmark, down from 9.2% for 3.7 Flash and beside the Claude line. That does not close the problem we described in our prompt-injection post, but it moves the arithmetic in the right direction for exactly the class of work the fast lane does.
§ 03Why day one was safe
We have a rule for model upgrades: switch on day one when the successor is a clean superset, and wait when it is not. Gemini 3.8 Flash is a clean superset. Same limits, same effort dial, same API surface, same price. We ran it on our infrastructure before the switch, and Gemini 3.7 Flash stays fully wired as a one-line rollback; Google announced no shutdown date for it.
We did the same with 3.7 Flash on the day Google shipped it, and with Fable 5.1 yesterday. The reasoning is the same each time: when waiting buys no safety, it only costs users capability.
§ 04The honest caveat
Google says the gains come from a design choice: 3.8 Flash “works harder,” running extra reasoning steps and calling tools iteratively on complex tasks, and “might use more tokens to maximize performance, especially at higher effort levels.” A bill is price times tokens. Same per-token price, possibly more tokens per job.
Two decisions follow from that sentence. We kept the medium thinking setting rather than raising it, because Google names medium as the level that still solves complex agentic tasks while limiting token use, and a higher setting would compound the increase on the cheap tier. And we are measuring, not estimating: Flash cost per turn and per shift, the week before the switch against the week after. If the number climbs more than we like, the levers are a lower thinking setting or 3.7 Flash, both one line away. Your pricing does not move in any of those cases.
§ 05What does not change
Your pricing. A full shift of real work still runs about $25, and the cost depends purely on how much work you assign.
Your setup. There is no model picker and no migration step. If a chat or an employee is on a Flash tier, it is on Gemini 3.8 Flash.
Your other tiers. Core, Max, and Creative run exactly what they ran before.
§ 06What we are watching
Three things over the next few weeks: Flash token use and cost per turn against the pre-switch baseline, any change in refusal behavior on the Gemini lane under Google’s stricter safety posture, and independent replication of Google’s benchmarks. Updates land on the launch post, and anything that changes CellCog’s routing lands here.
Q1Which CellCog tiers run on Gemini 3.8 Flash now?
Two: Agent Flash and Team Flash. Both moved from Gemini 3.7 Flash on September 2, 2026 with their operating point otherwise untouched, including the medium thinking setting we have run since August 13. Agent Core, Agent Max, Team Core, and Team Max run on Claude Fable 5.1; Agent Creative runs on Claude Opus 5. How the tiers fit together is in Flash, Core, Max.
Q2Why did you switch on day one?
Because Gemini 3.8 Flash is a drop-in successor. Google’s developer guide lists the same 1,048,576-token context window, the same 65,536-token output limit, and the same low, medium, and high thinking levels as 3.7 Flash, at the same price. We verified it on our infrastructure before switching, and 3.7 Flash remains one setting away as a fallback. We did the same with 3.7 Flash on August 13 and with Fable 5.1 on September 1: when the successor is a clean superset, waiting costs users capability for nothing.
Q3Will Flash work cost more now?
Your pricing is unchanged. On our side, the per-token price is identical to 3.7 Flash, but Google states that 3.8 Flash executes extra reasoning steps and tool calls on complex tasks and may use more tokens, especially at higher effort levels. That is why we kept the medium setting rather than raising it, and why we are measuring Flash cost per turn over the first seven days. If it climbs more than we like, the levers are a lower thinking setting or 3.7 Flash, both one line away.
Q4How much better is Gemini 3.8 Flash than 3.7 Flash?
On Google’s announcement table, every published row moved up: DeepSWE v1.1 73.7% vs 65.3%, OSWorld-2.0 59.0% vs 50.6%, Terminal-bench 2.1 89.4% vs 85.8%, Vals Finance Agent v2 61.4% vs 59.0%, HLE-Verified 54.9% vs 53.6%. Google also charts both 3.8 models at a 5.5% prompt-injection attack success rate on Gray Swan’s benchmark, down from 9.2% for 3.7 Flash. These are Google’s runs; independent replication is the open question, and we will note what holds up.
Q5What is the difference between Flash and Core now?
The model class. Flash tiers run a Flash-class model for fast, high-volume work where speed and cost matter more than maximum reasoning depth; Core and Max run Claude Fable 5.1 at a lighter and a deeper reasoning setting. Agent tiers are one agent working in a loop with you or running an AI employee’s shift; Team tiers are several agents challenging each other’s findings.
