Field intelligence for AI-first professionalsVol. II · Nº 61 · Thursday, September 24, 2026
← Phantom Notes
Nº 57AI ModelsSeptember 24, 202611 min read

Gemini 3.8 Flash Is the Third Flash in Six Weeks. Google's Pro Page Still Says 'Coming Soon'.

Google shipped Gemini 3.8 Flash on September 2, its third Flash model since July 21, at the same $0.75/$3.75 intro price as 3.7 Flash. The catch is in Google's own wording: the model spends more tokens per task by design, so cheaper per token is not cheaper per task. We checked the model card, the pricing page, the API docs, and Artificial Analysis, which re-versioned its index mid-month, and we looked hard at the Pro model Google keeps listing as coming soon.

By T.W. Ghost


Google shipped Gemini 3.8 Flash on September 2, 2026. That makes three Flash models in six weeks: 3.6 Flash on July 21, 3.7 Flash on August 13, and 3.8 Flash 20 days after that. The model card says 3.8 Flash is based on 3.7 Flash. Google is iterating one branch at high speed, and the price has not moved since August.

What has moved is the token bill. Google's launch post says the model "might use more tokens to maximize performance, especially at higher effort levels", and the API docs say the increase is by design. Cheaper per token is not the same as cheaper per task. That is the first story. The second is the one Google keeps not telling: the Pro model announced alongside 3.5 Flash at I/O in May is still labeled "3.5 Pro coming soon" on DeepMind's site as of today, September 24. We read the pricing page, the model card and its evaluation PDF, the API docs, and Artificial Analysis. Here is what held up.


Same Price, More Tokens, by Design

Gemini 3.8 Flash
AnnouncedSeptember 2, 2026, GA (no preview suffix)
Model IDgemini-3.8-flash
API price (intro)$0.75 in / $3.75 out per 1M, through Dec 31, 2026
API price (standard)$1.50 in / $7.50 out per 1M, from Jan 1, 2027
Context1,048,576 tokens in / 65,536 out
ModalitiesText, image, video, audio, PDF in; text out
Knowledge cutoffMarch 2026 (some domains only to Jan 2025, per the model card)
Thinking levelsLow, medium (default), high. Minimal is not supported
AA Intelligence Index41 (high) on the current v4.3.2 index; 3.7 Flash reads 39

The price is identical to 3.7 Flash, asterisk included. Google's pricing page lists the same introductory rate, the same December 31, 2026 expiry, and the same doubling to $1.50/$7.50 on January 1, 2027, with Batch and Flex at half ($0.375/$1.875) and caching on the same curve. The model-card footnote spells it out: "For 3.7 and 3.8 Flash, introductory price expires on December 31, 2026."

3.7 Flash is not deprecated. Google's blog says developers can "continue to rely on Gemini 3.7 Flash, which remains fully supported for efficiency-first workloads." It is still on the pricing page at the same rate and still Stable on the models page. "Efficiency-first workloads" is Google's polite way of saying the new model is not the efficient one.

Google says it will spend more tokens, and says so twice. The launch post describes the model on complex tasks as "executing extra reasoning steps, and calling tools iteratively", and adds that it "might use more tokens to maximize performance, especially at higher effort levels." The API docs are blunter: Gemini 3.8 Flash "can use more tokens on longer running and complex tasks, by design." Google publishes no percentage. Vellum estimates roughly 30% more output tokens and The Decoder roughly 40% more tokens per task. Those are press estimates, not Google's figures, so we print them labeled and nothing more.

API migration is small but real. Replace thinking_budget with the thinking_level string enum (low, medium, high), remove candidate_count, and stop sending temperature, top_p, and top_k, deprecated since July 21, 2026.


The Benchmarks, Sorted by Who Reported Them

House rules: independent numbers first, stated as fact; Google's numbers labeled as Google's.

First, a correction to our own record. Artificial Analysis re-versioned its Intelligence Index three times this month, most recently to v4.3.2 on September 19. On launch day, AA scored 3.8 Flash at 59 (high), up 3 from 3.7 Flash's 56, and called it "cheapest model at its level of intelligence." Those numbers, including the 56 we printed for 3.7 Flash in August, are on a retired scale. On the current v4.3.2 index, Gemini 3.8 Flash (high) scores 41, medium 40, low 33, and Gemini 3.7 Flash (high) reads 39. The gap survived the re-scoring; the absolute numbers did not.

Metric (Artificial Analysis, v4.3.2)Gemini 3.8 Flash
Intelligence Index41 (high), 40 (medium), 33 (low)
Output speed282.5 tokens/s
Time to first answer token15.64 s
Blended price (7:2:1 cache:input:output)$0.58 per 1M
Cost per Intelligence Index task$1.24 (high), $0.93 (medium)
Terminal-Bench 4.020%

Two things to notice. The 15.64 second time to first token is worse than the 10.90 seconds AA currently shows for 3.7 Flash, which fits a model that thinks longer before answering; both are wrong for a live chat box and fine for a pipeline. And AA's "cheapest at its level of intelligence" line, true on September 2, has been overtaken: GPT-6 Sol (max) now scores 48 at $1.06 per index task, smarter and cheaper per task than 3.8 Flash's 41 at $1.24. Per token, 3.8 Flash still wins ($0.58 blended against Sol's $1.54). That is the split Google's "by design" language predicts.

One more independent data point: Datacurve's public DeepSWE v1.1 leaderboard (113 tasks, updated September 22) has Claude Opus 5 at 74% plus or minus 4 and Gemini 3.8 Flash at 74% plus or minus 1. On that one long-horizon coding benchmark, a Flash model is statistically tied with Opus 5.

Now the vendor table. These are Google's self-reported numbers from the DeepMind evaluation PDF: Gemini rows are pass@1 (most self-computed by Google, the DeepSWE run at high thinking; Terminal-bench 4.0, GDPval-AA v2, and the two Vals rows pulled from those leaderboards), non-Gemini rows "sourced from providers' self reported numbers unless otherwise mentioned". We show 3.8 Flash, its predecessor, and the strongest rival column, Claude Opus 5.

Benchmark (Google's table)3.8 Flash3.7 FlashClaude Opus 5
DeepSWE v1.173.7%65.3%74.0%
GDPval-AA v2 (Elo)154514821824
Vals Finance Agent v261.4%59.0%58.6%
Harvey Legal Agent (all pass)10.0%8.8%6.7%
Terminal-bench 2.189.4%85.8%89.1%
Terminal-bench 4.019.1%11.2%51.8%
GDP.PDF (all pass)35.0%34.0%37.0%
CharXiv Reasoning86.2%84.5%83.7%
LVBench (long video)87.8% agentic / 87.1% static85.4%75.4%
HLE-Verified54.9%53.6%54.4%
OSWorld-2.0 (partial, batch tool)59.0%50.6%75.4%
BioMysteryBench, Human Solvable88.8%87.1%90.1%
BioMysteryBench, Human Difficult56.5%43.5%49.4%
LABBench286.2%82.1%84.2%

Google's summary holds: 3.8 Flash beats 3.7 Flash on all 14 rows, with no regressions, and is best of the six-model table on 8. Against Claude Opus 5 it is 8 wins and 6 losses.

Three honest observations Google's marketing does not lead with:

  • The losses are where the agent work is. The three biggest gaps to Opus 5 are all agentic: Terminal-bench 4.0 (19.1% vs 51.8%), GDPval-AA v2 (1545 vs 1824 Elo), and OSWorld-2.0 (59.0% vs 75.4%). The wins cluster in science, video, and two Vals-sourced finance and legal benchmarks. That is a specialist, not a frontier model.
  • Terminal-bench 2.1 is 89.4%, not 90.8%. Two outlets printed 90.8% for 3.8 Flash and 81.6% for 3.7 Flash, attributing both to a Google Cloud developer guide that did not render for us. Google's model card and evaluation PDF say 89.4% and 85.8%. We print the model card and note the discrepancy, because that is the whole job.
  • Three harnesses, one table. The 3.7 Flash OSWorld-2.0 baseline here is 50.6%; its own model card said 47.9%. Google re-ran it with a batched tool call, pulled the Opus 5 row from Anthropic's Fable 5.1 blog post, and used "maximum thinking settings when available" for Terra and Sonnet 5 while its own DeepSWE run used high thinking.

The Prompt Injection Number, and Where It Came From

Google's blog says the 3.8 models "have also made a significant leap" in resisting prompt injection "as measured by Gray Swan" and stops there. The number everyone printed, a 5.5% attack success rate on the Gray Swan IPI benchmark, is not in the model card or the evaluation PDF. It comes from a chart in Google's launch materials, reproduced by Vellum and quoted by The Decoder, which also puts Claude Opus 5 at 4.8%, Gemini 3.7 Flash at 9.2%, and GPT-5.6 Sol at 27.0% (lower is better). It is a vendor chart; treat it as one. We are not printing an attempt count, because no source we read states one.

On the Frontier Safety Framework, the model card is direct: "Gemini 3.8 Flash does not have meaningful new capabilities or material increases in performance with respect to the domains outlined in our Frontier Safety Framework compared to Gemini 3.7 Flash." Same CBRN and cyber safeguards as last month, no new gating on the standard model.


Flash Cyber and the Fairwind Program

Gemini 3.8 Flash Cyber is the same model with, in Google's words, "a more permissive set of mitigations for cybersecurity," and it "is only available to trusted defenders who require a more comprehensive set of cyber capabilities." It replaces the 3.5 Flash Cyber pilot from July 21. Access runs through the Fairwind Program, announced the same day by Four Flynn, Google's VP of Security and Privacy.

Who qualifies. Google's blog lists "trusted government authorities, as well as critical infrastructure operators and software maintainers"; the Fairwind page claims "over 650 partners globally." Partners cannot "share, redistribute, or sell access", must use phishing-resistant MFA, and go through background checks. Permitted uses are "authorized threat simulation, reverse engineering, and malware analysis for defensive and academic research purposes"; "malicious tasks such as creating malware are not permitted."

Google's numbers for the Cyber variant (vendor, DeepMind Cyber page): CyberGym pass@1 of 86.2% against 85.6% for GPT-5.5-Cyber, and CWE-Bench patching at 47.2% pass@1 against 47.8% for what the page labels as Claude Fable 5, which Google's blog frames as near parity "at a significantly lower cost." None of it is independently reproduced, and unless you are a government, an infrastructure operator, or a core platform, none of it is available to you anyway.


The Pro Question, Answered Only With What Google Has Said

We are going to be careful here, because the internet is not.

What Google has said on record. On July 21, in the 3.6 Flash launch post, Google wrote: "Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it's ready." The same post said: "We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress." A Google spokesperson told Fortune on August 10 that 3.5 Pro was in testing. DeepMind's Pro page, which we loaded today, September 24, still carries the label "3.5 Pro coming soon" next to Gemini 3.1 Pro. Google has never said 3.5 Pro is dead.

What the press has said. Fortune reported in August that Google had expected 3.5 Pro in June and again in mid-July and missed both, and in September cited the Wall Street Journal's reporting that internal candidates were discarded because they did not improve enough over Flash. The "shelved" claim traces to a SemiAnalysis institutional note from around August 10, relayed by press outlets that also noted Google had issued no statement. Fortune also reported the reorganization: Demis Hassabis moved to chairman on August 7 and Koray Kavukcuoglu, formerly CTO, took over day-to-day at DeepMind.

What Kavukcuoglu said this week. At The Information's AI Agenda Live summit on September 23 and 24, Kavukcuoglu said Google "took a little bit of a step back" on 3.5 Pro, and that Gemini 4 is in an early phase of post-training with an intent to ship as soon as possible. He did not confirm the model is gone. He did not give a date.

What exists on the price list. The newest shipped Pro is still Gemini 3.1 Pro Preview, released February 19, 2026, at $2.00 in / $12.00 out per 1M for prompts up to 200K tokens and $4.00/$18.00 above that. No 3.5 Pro or Gemini 4 row exists on the pricing page, the models page, or the changelog.

So the accurate sentence is this: seven months after its last Pro model, Google has shipped three Flash models, kept a "coming soon" label on the Pro page, and put the next-generation model in early post-training with no date. Whether that is a Pro that arrives late or a Pro that gets skipped for Gemini 4, only Google can say, and it has chosen not to. If you are comparing frontier options today, look at Claude Opus 5.5, GPT-6 Astra, and Grok 4.7, which is where the current leaderboard actually lives.


Where You Can Run It

The Gemini API in Google AI Studio and Android Studio, the Gemini Enterprise Agent Platform (formerly Vertex AI), Gemini Enterprise, Google Antigravity, and Stitch. Consumer-side, the Gemini app for Google AI Pro and Ultra subscribers, AI Mode in Search, and Gemini in Google Sheets. Free-tier Gemini app users do not get it as a selectable model; the API free tier does cover it. Computer use is in preview; image generation, audio generation, and the Live API are not supported on this model ID.

The rest of the September cadence:

  • September 15: gemini-3.8-live and gemini-3.8-live-extended-thinking GA.
  • September 18: Gemini 2.5 restricted to existing users; new projects are steered to 3.5 Flash-Lite or 3.8 Flash.
  • September 22: gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts GA.

Three weeks, four more 3.8-branded model IDs, zero Pro.


Verified vs Unconfirmed: The Scorecard

ClaimVerdict
Released Sep 2, 2026; third Flash GA since Jul 21; ID is gemini-3.8-flashVerified (Gemini API changelog, models page)
Intro price $0.75/$3.75 through Dec 31, 2026, then $1.50/$7.50; identical to 3.7 FlashVerified (Google's pricing page, model-card footnote)
"30 to 40% more tokens per task"Unverified (Vellum and The Decoder estimates; Google publishes no percentage)
AA Index 41 vs 39 for 3.7 Flash; 282.5 tok/s; 15.64 s TTFT; $0.58 blendedVerified (Artificial Analysis, v4.3.2, read Sep 24)
Terminal-bench 2.1 at 90.8%Contradicted (model card and evaluation PDF say 89.4%; the 90.8% is attributed to a Google Cloud guide we could not load)
Gray Swan IPI 5.5% attack successVerified as a vendor chart (reproduced by Vellum, quoted by The Decoder; not in the model card or PDF)
Gemini 3.5 Pro shelvedUnverified (SemiAnalysis note via press; Google says "coming soon" as of Sep 24)

Who Should Use It

Use Gemini 3.8 Flash if you already run 3.7 Flash agents on tasks where a few more reasoning steps pay for themselves: long-horizon coding (it ties Opus 5 on Datacurve's independent DeepSWE board), document-heavy finance and legal pipelines, video understanding, and scientific literature work. Same price as 3.7 Flash until December 31, so the upgrade costs only the extra tokens. Budget against $1.50/$7.50, and run it at medium unless high provably moves your metric; AA's cost per task drops from $1.24 to $0.93 for one index point.

Skip it for anything latency-sensitive (15.64 seconds to first token), for high-volume classification where Google itself points you at 3.7 Flash, and for general computer-use and terminal agents, where Google's own table shows 19.1% on Terminal-bench 4.0 against 51.8% for Opus 5. If per-task cost is your metric, GPT-6 Sol is both smarter and cheaper on the current index.

The frontier is elsewhere: Claude Opus 5.5 (max) leads the v4.3.2 index at 58, with Claude Fable 5.1 (max) and GPT-6 Astra (max) at 53 and Grok 4.7 (xhigh) at 46. Gemini 3.8 Flash at 41 is the best Google has shipped, and it is not a frontier model, which is precisely why the Pro page matters. Our model comparison has the current standings side by side.


Sources


*Not sure whether a cheap, deliberate agent engine fits how you work, or whether you need the frontier Google has not shipped? Take the free 2-minute quiz and get matched. Then read the other model that landed this week with fine print of its own: Grok 4.7.*

Which model should you be using?

Three minutes, twelve questions, one defensible answer.

Take the quiz →