Google released Gemini 3.8 Flash on September 2, 2026, alongside a restricted cybersecurity variant, Gemini 3.8 Flash Cyber. By Google’s own count it is the third Flash release in six weeks, and the headline is the thing that did not move: the price. 3.8 Flash ships at the same introductory rate as 3.7 Flash, $0.75 per million input tokens and $3.75 per million output tokens, with the same 1M-token context window and the same three thinking levels. What moved is the benchmark column, most of all on agentic coding and professional work, where a Flash-priced model now sits within a point of Claude Opus 5 on several rows.
This page is built the way our Fable 5.1 record and our 3.7 Flash post were: from Google’s own pages, read the day they went up, with vendor-run numbers labeled as such.
On this page · 7 sectionsOpen
- Google released Gemini 3.8 Flash on September 2, 2026, three weeks after 3.7 Flash, as a generally available model with the id gemini-3.8-flash, a 1,048,576-token context window, a 65,536-token output limit, and three thinking levels (low, medium, high; default medium).
- The price did not move: $0.75 per million input tokens and $3.75 per million output tokens, the same introductory rate as 3.7 Flash, through December 31, 2026. From January 1, 2027 the standard rate is $1.50 and $7.50.
- Google’s own comparison table puts 3.8 Flash ahead of 3.7 Flash on every row it publishes, led by DeepSWE v1.1 (73.7% vs 65.3%), OSWorld-2.0 (59.0% vs 50.6%), and BioMysteryBench’s hard tier (56.5% vs 43.5%).
- Against frontier models, 3.8 Flash ties Claude Opus 5 on DeepSWE v1.1 (73.7% vs 74.0%) and Terminal-bench 2.1 (89.4% vs 89.1%) and leads every listed model on Vals Finance Agent v2 and Harvey’s legal benchmark, at $3.75 per million output tokens against Opus 5’s $25. It trails badly on Terminal-bench 4.0 (19.1% vs 51.8%) and OSWorld-2.0 (59.0% vs 75.4%).
- The caveat Google states itself: 3.8 Flash ‘works harder’, running extra reasoning steps and tool calls, and may use more tokens at higher effort levels. Same unit price, possibly bigger bills; Google’s own advice for cost-first workloads is lower effort or staying on 3.7 Flash, which remains fully supported.
- Gemini 3.8 Flash Cyber is a second variant on the same core, restricted to vetted defenders through Google’s new Fairwind Program: 86.2% on CyberGym and 47.2% pass@1 on CWE-Bench against Claude Fable 5’s 47.8%. Both 3.8 models also post the lowest prompt-injection attack success rates in Google’s Gray Swan chart apart from Claude Opus 5.
- For CellCog users: the Flash tiers, Agent Flash and Team Flash, moved to Gemini 3.8 Flash the same day, at the same thinking setting. Your pricing does not change; the operator’s side is in our day-one update.
§ 01What shipped
Google’s announcement names two variants powered by one foundational model. The standard model is the one most readers can use; the developer guide on Google Cloud carries the specification table.
| Item | Detail |
|---|---|
| Model id | gemini-3.8-flash |
| Launch stage | Generally available |
| Context window | 1,048,576 tokens |
| Maximum output | 65,536 tokens |
| Inputs and output | Text, image, audio, video in; text out |
| Thinking levels | Low, medium, high (default medium) |
| Introductory price | $0.75 input / $3.75 output per 1M tokens, thinking tokens billed as output, through December 31, 2026 |
| Standard price | $1.50 input / $7.50 output per 1M tokens from January 1, 2027 |
| Developer access | Gemini API via Google AI Studio, Android Studio, Antigravity, Stitch; Gemini Enterprise for businesses |
| Consumer access | Gemini app, AI Mode in Google Search, Gemini in Sheets, for Google AI Pro and Ultra subscribers |
| Gemini 3.7 Flash | Remains fully supported for efficiency-first workloads; no shutdown announced |
On paper this is a pure drop-in: the same limits, the same effort dial, the same API surface, a new model id. That is exactly the shape 3.7 Flash had against 3.6 Flash three weeks ago.
§ 02The benchmarks: Google’s full table
The announcement publishes a six-model comparison, and it is unusually generous about naming the competition: Claude Opus 5, Claude Sonnet 5, GPT-5.6 Sol, and GPT-5.6 Terra sit beside both Flash generations. These are Google’s runs, under Google’s published evaluation methodology; independent replication is pending, as it always is on release day. Prices are as Google listed them, per million tokens without caching.
| Benchmark | Gemini 3.8 Flash | Gemini 3.7 Flash | Claude Opus 5 | Claude Sonnet 5 | GPT-5.6 Sol | GPT-5.6 Terra |
|---|---|---|---|---|---|---|
| DeepSWE v1.1 (long-horizon software engineering) | 73.7% | 65.3% | 74.0% | 53.8% | 72.7% | 69.6% |
| GDPVal-AA v2 (knowledge work, Elo) | 1545 | 1482 | 1824 | 1584 | 1710 | 1528 |
| Vals Finance Agent v2 | 61.4% | 59.0% | 58.6% | 53.9% | 53.8% | 54.4% |
| Harvey’s Legal Agent Benchmark (all pass rate) | 10.0% | 8.8% | 6.7% | 5.0% | 2.5% | 0.8% |
| Terminal-bench 2.1 (agentic terminal coding) | 89.4% | 85.8% | 89.1% | 80.4% | 88.8% | 87.4% |
| Terminal-bench 4.0 (general agent capabilities) | 19.1% | 11.2% | 51.8% | 12.4% | 37.3% | 23.6% |
| GDP.PDF (expert document comprehension, all pass rate) | 35.0% | 34.0% | 37.0% | 28.0% | 40.0% | 29.0% |
| CharXiv Reasoning (no tools) | 86.2% | 84.5% | 83.7% | 70.1% | 85.8% | 85.9% |
| LVBench (long video understanding) | 87.8% agentic / 87.1% static | 85.4% | 75.4% | 68.5% | 82.1% | 78.9% |
| HLE-Verified (expert reasoning) | 54.9% | 53.6% | 54.4% | 31.0% | 54.5% | 51.1% |
| OSWorld-2.0 (agentic computer use, partial score) | 59.0% | 50.6% | 75.4% | 42.6% | 62.6% | 50.2% |
| BioMysteryBench, human solvable | 88.8% | 87.1% | 90.1% | 87.5% | 79.5% | 83.8% |
| BioMysteryBench, human difficult | 56.5% | 43.5% | 49.4% | 34.1% | 44.7% | 49.4% |
| LABBench2 (biology research tasks) | 86.2% | 82.1% | 84.2% | 80.1% | 82.1% | 81.2% |
Two readings, one chart each.
First, the generation jump. Every row moved up from 3.7 Flash, and the biggest moves are on the agent-shaped rows: DeepSWE, OSWorld, Terminal-bench 4.0, and the hard tier of the biology benchmark.
Second, the frontier-at-Flash-price test. Put 3.8 Flash next to the two most expensive models in the table and the story splits cleanly in two. On agentic coding, finance, legal, and expert reasoning, the $3.75 model is level with or ahead of the $25 and $20 models. On long-horizon general agency and computer use, it is not close.
The rows where 3.8 Flash wins outright are the professional-task benchmarks: Vals Finance Agent v2 (61.4%, ahead of Opus 5’s 58.6%) and Harvey’s Legal Agent Benchmark (10.0%, ahead of Opus 5’s 6.7% and far ahead of GPT-5.6 Sol’s 2.5%). Those are the two Google’s announcement leads with, and they are the rows a business buyer reads first. Google’s cost-per-task chart for DeepSWE, using Datacurve AI’s data, makes the same point visually: 3.8 Flash lands within 0.3 points of Opus 5 at a fraction of the per-task cost.
The rows where it loses are just as instructive. Terminal-bench 4.0 is the general-agent successor to the coding benchmark 3.8 Flash aces, and the gap there (19.1% against Opus 5’s 51.8%) is the widest in the table. OSWorld-2.0, agentic computer use, sits 16 points behind Opus 5. GDPVal-AA, a knowledge-work Elo, has 3.8 Flash at 1545 against Opus 5’s 1824 and Sol’s 1710. Read together: a Flash-class model is now a frontier model on bounded, well-defined agent tasks, and still a Flash-class model on open-ended, long-horizon ones. That is a useful line to know when you decide which work to route where.
A note on the developer guide. Google’s Cloud developer guide publishes a second, shorter 3.8-vs-3.7 table with four rows the announcement does not carry.
| Benchmark | 3.8 Flash | 3.7 Flash |
|---|---|---|
| SWE-Bench Pro | 61.6% | 60.4% |
| SWE-Atlas | 51.9% | 48.0% |
| τ³-bench Banking | 38.1% | 30.9% |
| Humanity’s Last Exam (unverified set) | 45.4% | 45.7% |
One row disagrees between the two Google pages: the announcement table lists Terminal-bench 2.1 at 89.4% vs 85.8%, while the developer guide lists it at 90.8% vs 81.6%. Both are Google’s numbers, presumably from different run configurations. We show the announcement’s figures above and note the guide’s here; we have not averaged them. Also worth separating: the 54.9% HLE-Verified figure in the announcement and the 45.4% figure in the guide are different test sets, not a nine-point jump on one exam.
§ 03The pricing, and the token caveat
The unit prices are unchanged from 3.7 Flash, including the same cliff on New Year’s Day. Google’s pricing page and the announcement’s own footnote agree. Against the competition in Google’s table, the output price is the whole argument.
| Model | Input | Output |
|---|---|---|
| Gemini 3.8 Flash | $0.75 ($1.50 from Jan 1, 2027) | $3.75 ($7.50 from Jan 1, 2027) |
| Gemini 3.7 Flash | $0.75 ($1.50 from Jan 1, 2027) | $3.75 ($7.50 from Jan 1, 2027) |
| Claude Sonnet 5 | $2.00 | $10.00 |
| GPT-5.6 Terra | $2.00 | $12.00 |
| GPT-5.6 Sol | $4.00 | $20.00 |
| Claude Opus 5 | $5.00 | $25.00 |
The introductory rate is a real deadline, not a launch flourish: the same footnote that sets $0.75 and $3.75 sets $1.50 and $7.50 for January 1, 2027.
| Token type | Introductory (through Dec 31, 2026) | Standard (from Jan 1, 2027) |
|---|---|---|
| Input | $0.75 | $1.50 |
| Output, including thinking tokens | $3.75 | $7.50 |
The part that matters more than either table is a sentence in the announcement. Google says the performance gains “stem from a core design choice: 3.8 Flash works harder,” executing extra reasoning steps and calling tools iteratively on complex tasks, and that “at times, the model might use more tokens to maximize performance, especially at higher effort levels.” A bill is price times tokens. Two models at an identical per-token price can produce different bills for the same job, and Google is telling you in advance that this one leans toward more tokens.
Google’s own remedy is in the same paragraph: for applications where compute efficiency is the primary constraint, use a lower effort level, or stay on 3.7 Flash, which remains fully supported. If you run agents at volume, the right move on day one is a measurement, not a migration: run your real workload at your real effort level on both model ids and compare token counts, not just quality.
§ 04Gemini 3.8 Flash Cyber: a model for defenders only
The second variant shares the core and is tuned for cybersecurity, with a more permissive set of cyber mitigations than the standard model. That is why Google gates it: access goes through the new Fairwind Program, open to trusted government authorities, critical-infrastructure operators, and software maintainers, by application. Google’s framing is deliberate about direction: it says it prioritized vulnerability fixing over offensive capabilities like exploitation.
| Benchmark | 3.8 Flash Cyber | Against |
|---|---|---|
| CyberGym pass@1, vulnerability discovery in C/C++ | 86.2% | GPT-5.5-Cyber 85.6%, Mythos 5 83.8%, GPT-5.6 Sol 83.6%, Gemini 3.5 Flash Cyber 77.5% |
| Internal vulnerability discovery, 20 languages | 71.0% | Gemini 3.7 Flash 58.9%, Gemini 3.5 Flash Cyber 46.6% |
| CWE-Bench pass@1, automated patching (run by Collinear) | 47.2% | Claude Fable 5 47.8%, at more than twice the cost per rollout on Google’s chart; GPT-5.6 Sol and Gemini 3.7 Flash in the mid-40s |
| Chrome Security team, correct vulnerability patches in Chrome | 2.6x more | The best commercial models, which are much larger |
| Wiz internal penetration-testing benchmark | 7.5 to 9.7 points higher recall | Other leading frontier models, at 2.3 to 5.2x lower cost |
Google also says its Cloud Vulnerability Research team used the model to find a critical foundational vulnerability in under two hours, work it describes as usually taking months. The standard 3.8 Flash ships with safeguards against misuse in CBRN and cyber-offense domains under Google’s Frontier Safety Framework; the Cyber variant relaxes the cyber side, which is the whole reason it is not on the public menu.
The prompt-injection chart
The result in this release most relevant to anyone who runs agents with real access is not a cyber-variant number at all. Google publishes a Gray Swan indirect-prompt-injection chart for sixteen models, attack success rate within 15 attempts, lower is better. Both 3.8 models land near the bottom, beside Anthropic’s Claude line and a long way from the open-weight and GPT-5.6 entries.
| Model | ASR@15 |
|---|---|
| DeepSeek V4 Pro | 60.1% |
| Kimi K3 | 52.7% |
| Grok 4.6 | 51.8% |
| GPT-5.6 Luna | 50.0% |
| GPT-5.6 Terra | 37.3% |
| GLM 5.3 | 31.5% |
| Qwen 3.8 | 28.6% |
| GPT-5.6 Sol | 27.0% |
| Muse Spark 1.2 | 24.2% |
| Gemini 3.7 Flash | 9.2% |
| Claude Opus 4.8 | 8.0% |
| Claude Sonnet 5 | 6.7% |
| Claude Fable 5 | 6.5% |
| Gemini 3.8 Flash Cyber | 6.0% |
| Gemini 3.8 Flash | 5.5% |
| Claude Opus 5 | 4.8% |
The move from 9.2% on 3.7 Flash to 5.5% on 3.8 Flash is the kind of number that decides whether a cheap model is allowed to read untrusted email and web pages on your behalf. We have written about why prompt injection is a different problem for persistent AI workers than for a chatbot; a model that fails one attack in eighteen instead of one in eleven does not close that problem, but it changes the arithmetic for anything you route to the fast lane.
§ 05Who should care, and what to do
If you build on Gemini 3.7 Flash: the migration is a model-id change. Same context window, same output limit, same three thinking levels. Re-run your evals, and re-read your token usage at the effort level you actually use before you flip production, because Google has told you the token count may move.
If you are choosing an AI agent or AI employee platform: note the cadence. Three Flash releases in six weeks from Google, Fable 5.1 from Anthropic yesterday, Qwen3.8-Max-0902 overnight. The model layer now rotates monthly or faster. The question to ask a platform is not which model it runs today but whether the work you build on it survives the next swap: the roles, the memory, the accumulated context, the tasks in flight.
§ 06What it means for CellCog users
CellCog’s Flash tiers, Agent Flash and Team Flash, moved to Gemini 3.8 Flash today, the same way the fast lane moved to 3.7 Flash the day Google shipped it. Same context window, same output limit, the same medium thinking setting, same price, and a better prompt-injection number for the class of work a fast lane does. The one new variable is Google’s own token warning, and we are measuring it rather than guessing: Flash cost per turn, the week before against the week after. The operator’s side of the switch, and what we are watching, is in the day-one update. How the tiers fit together is in Flash, Core, Max.
Your pricing is not tied to any of this. A full shift of real work runs about $25, and the cost depends purely on how much work you assign.
§ 07The record
As of September 2, 2026: Gemini 3.8 Flash is generally available at the same introductory price as 3.7 Flash, with Google’s six-model comparison table published; Gemini 3.8 Flash Cyber is available by application through the Fairwind Program; Gemini 3.7 Flash remains supported with no shutdown date. If independent replications land materially different from Google’s numbers, if Google reconciles the two Terminal-bench 2.1 figures, or if the 3.7 Flash deprecation schedule changes, the update happens here, same day.
Q1When was Gemini 3.8 Flash released?
September 2, 2026. Google’s announcement calls it the third Flash release in six weeks, following 3.7 Flash on August 13. It is generally available, not a preview, with the model id gemini-3.8-flash.
Q2What are the confirmed specifications?
From Google’s developer guide: 1,048,576-token context window, 65,536-token maximum output, text, image, audio, and video input with text output, and three thinking levels, low, medium, and high, with medium as the default. Those are the same limits and levels as Gemini 3.7 Flash.
Q3How does Gemini 3.8 Flash compare to Claude Opus 5 and GPT-5.6?
On Google’s own table it is within a point of Claude Opus 5 on DeepSWE v1.1 (73.7% vs 74.0%) and Terminal-bench 2.1 (89.4% vs 89.1%), ahead of every listed model on Vals Finance Agent v2 (61.4%) and Harvey’s Legal Agent Benchmark (10.0%), and level on HLE-Verified (54.9% vs Opus 5’s 54.4% and GPT-5.6 Sol’s 54.5%). The gaps that remain are long-horizon general agency and computer use: Terminal-bench 4.0 19.1% vs 51.8% for Opus 5, OSWorld-2.0 59.0% vs 75.4%, and a GDPVal-AA knowledge-work Elo of 1545 vs 1824. Opus 5 costs $25 per million output tokens to 3.8 Flash’s $3.75.
Q4Why might 3.8 Flash cost more than 3.7 Flash if the prices are identical?
Because the bill is price times tokens, and Google states that 3.8 Flash executes extra reasoning steps and iterative tool calls on complex tasks, especially at higher effort levels. Two models at the same per-token price can produce different bills for the same job. Google says 3.7 Flash remains fully supported for efficiency-first workloads, with no shutdown announced. If you run agents at volume, measure token usage at your effort level before switching.
Q5What is Gemini 3.8 Flash Cyber?
A variant of the same model tuned for cybersecurity defense, with a more permissive set of cyber mitigations, which is why it is gated behind Google’s Fairwind Program. Google reports 86.2% pass@1 on CyberGym, 71.0% on an internal 20-language vulnerability-discovery benchmark, 47.2% pass@1 on the CWE-Bench patching benchmark against Claude Fable 5’s 47.8%, 2.6 times more correct Chrome vulnerability patches than the best larger commercial models, and 7.5 to 9.7 points higher recall on Wiz’s internal penetration-testing benchmark at 2.3 to 5.2 times lower cost.
Q6Does CellCog run on Gemini 3.8 Flash?
Yes, since September 2, 2026. CellCog’s Agent Flash and Team Flash tiers moved from Gemini 3.7 Flash to 3.8 Flash the day Google released it, with the same medium thinking setting and the same limits. Core, Max, and Creative tiers are unchanged. Your pricing does not change: a full shift of real work runs about $25, and the cost depends purely on how much work you assign. The operator’s side, including what we are measuring, is in our day-one update.
