Back to Phantom Notes
AI Models

Gemini 3.6 Flash Is Not Smarter Than 3.5. Google Says That Is the Point.

July 21, 202610 min readBy T.W. Ghost
GoogleGeminiGemini 3.6 FlashGemini 3.5 Flash-LiteFlash CyberAI ModelsBenchmarksLLM PricingAgentic AI

Two months ago, when Gemini 3.5 Flash launched, we wrote that "Flash" no longer means cheap. Today Google answered with a price cut, and something stranger: a new flagship Flash model that is not any smarter than the old one.

Gemini 3.6 Flash scores exactly 50 on the Artificial Analysis Intelligence Index. So does Gemini 3.5 Flash. Same score, same rank tier, no intelligence gain at all. And Google is not hiding this. The entire launch pitch is built around what changed instead: the model does the same work in half the time, with fewer tokens, at a lower output price.

That is a genuinely unusual thing for an AI lab to ship in 2026, and whether it matters to you depends entirely on what you use models for. Here is the full picture, with the vendor numbers labeled as vendor numbers.


What Actually Launched

Three models, announced July 21, 2026. Two you can use today. One you cannot use at all.

Gemini 3.6 FlashGemini 3.5 Flash-Lite
API price$1.50 in / $7.50 out per 1M$0.30 in / $2.50 out per 1M
Cached input$0.15 per 1M (90% off)Cached rates available
Context1M tokens in / 64K out1M tokens in / 64K out
Knowledge cutoffMarch 2026March 2026
ModalitiesText, image, audio, video, PDF in; text outSame
Output speed303.6 tok/s (AA measured, #1 of 186)~350 tok/s (AA measured)

The third model, Gemini 3.5 Flash Cyber, is a cybersecurity specialist available exclusively to governments and trusted partners. More on that below.

Both usable models are live now in the Gemini API through Google AI Studio and Android Studio, and in the consumer Gemini app. Flash-Lite is also rolling out inside Google Search, and 3.6 Flash powers Google Antigravity.

One pricing note before you celebrate the cut: only the output price dropped, from $9.00 to $7.50 per 1M tokens. Input is unchanged at $1.50. There are also Batch and Flex tiers at half price ($0.75/$3.75), a Priority tier at nearly double ($2.70/$13.50), and context caching adds a storage fee of $1.00 per 1M tokens per hour. Read your own bill, not the headline.


The Real Story Is Efficiency, and It Is Independently Verified

Google's launch claim is that 3.6 Flash lowers the cost per agentic task, not that it thinks better. Unusually, the independent numbers back the claim.

Artificial Analysis, which runs its own benchmarks rather than reprinting vendor decks, measured:

  • 17% fewer output tokens to complete its Intelligence Index compared to 3.5 Flash
  • Cost per task down roughly 18%, from $0.59 to $0.50
  • Time per task cut in half, from 2.7 minutes to 1.3 minutes

That last number is the one that matters for agent workloads. If you run coding agents, browser agents, or long document pipelines, the model finishing in half the wall-clock time at the same quality is a bigger deal than two points on a benchmark.

Google's own numbers go further. The company reports up to 65% token savings on DeepSWE long-horizon engineering tasks. That is a vendor best-case figure and nobody has independently replicated it yet, so treat it as a ceiling, not an expectation.


The Google-Reported Benchmark Gains

All of the following are Google's launch numbers, not independent measurements. They are consistent with the third-party efficiency data, but they are the vendor grading its own homework:

BenchmarkGemini 3.6 FlashGemini 3.5 Flash
DeepSWE (long-horizon engineering)49%37%
MLE-Bench (ML engineering)63.9%49.7%
OSWorld-Verified (computer use)83.0%78.4%
GDPval-AA v2 (knowledge work, Elo)14211349

The pattern is coherent: every gain is agentic. Coding agents, computer use, multi-step knowledge work. Nothing here claims better raw reasoning, because there is not any.


The Asterisks

This site has a house rule: every launch gets its fine print read aloud. Here is the fine print.

Get the Weekly IT + AI Roundup

What changed this week in NinjaOne, ServiceNow, CrowdStrike, and AI. One email, every Monday.

No spam, unsubscribe anytime. Privacy Policy

Intelligence is flat. AA Index 50, ranked #21 of 186 models tracked. That is well above the average for its price tier (31), but it is exactly where 3.5 Flash already was. If your workload needs a smarter model, this launch gives you nothing.

The latency problem nobody is leading with. Gemini 3.6 Flash takes 11.54 seconds to produce its first token, per Artificial Analysis. The median for comparable reasoning models is 2.76 seconds. Once it starts, it is the fastest model in the world at 303.6 tokens per second. But that first-token wait makes it a poor fit for anything conversational. A user staring at a chat window for 11 seconds assumes the app is broken. A background agent does not care. Pick accordingly.

The comparison that stings. Gemini 3.1 Pro Preview costs $2.00/$12.00 at standard context. 3.6 Flash at $1.50/$7.50 now does most of the agentic work of a Pro-class model from seven months ago at roughly 60% of the price. Great if you are buying. Awkward if you built on 3.1 Pro.


Flash-Lite: The Quiet Overachiever

The second launch deserves more attention than it will get. Gemini 3.5 Flash-Lite at $0.30 input and $2.50 output is the budget play, and the numbers are startling for the price.

Google reports Terminal-Bench 2.1 at 54%, up from 31% for the previous 3.1 Flash-Lite. More interesting: Google's own charts show Flash-Lite beating the larger, more expensive Gemini 3 Flash on SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%). A Lite model outscoring the bigger model from one generation back, on coding and computer use, at a fraction of the cost.

At roughly 350 tokens per second as measured by Artificial Analysis, this is the model for high-volume, latency-sensitive work: agentic search, document processing, classification, extraction pipelines. It is also quietly rolling out inside Google Search itself, which tells you how much Google trusts it at scale.

If your workload was on the old Flash because it was cheap, Flash-Lite is probably your new home, and it is a better model than the one you were using.


Flash Cyber: The One You Cannot Have

Gemini 3.5 Flash Cyber is a cybersecurity specialist model, and the announcement is worth reading carefully because of what it does not say. There is no public API. There is no pricing. Access is, in Google's words, exclusively for governments and trusted partners through a limited-access pilot of CodeMender, Google's AI security-patching agent, "expanding over time."

Google claims frontier-competitive performance on CyberGym, a benchmark of finding and fixing real vulnerabilities. The interesting detail: that performance comes from CodeMender orchestrating multiple Flash Cyber agents per report, not from single-shot prompting. Small, cheap, specialized models running in swarms is a preview of where security tooling is going.

For readers of this site, the takeaway is one sentence: this model is news, not a product. You cannot select it, price it, or build on it. We will revisit if that changes.


The Roadmap, and a Correction

In our May coverage we passed along the expectation that Gemini 3.5 Pro would land within a month. It did not, and it still has not. Google says 3.5 Pro is "testing with partners" with no public date. We were early, and you should weight future "next month" roadmap claims, including Google's, accordingly.

The more interesting roadmap note buried in today's announcement: Google has started pre-training Gemini 4. The 3.x generation is now in its efficiency-harvesting phase while the next big training run cooks. That is usually what the end of a model generation looks like.


Who Should Use What

The buyer's guide, by workload:

  • Choose Gemini 3.6 Flash for coding agents, computer-use agents, and long-horizon pipelines where cost per completed task beats peak intelligence. The 1M context, top-ranked throughput, and cut output price are built for exactly this. Do not put it behind a chat box; the 11.5 second first token will make your product feel broken.
  • Choose Gemini 3.5 Flash-Lite for high-volume budget work: agentic search, document processing, classification, extraction. At $0.30/$2.50 it beats models that cost several times more, including Google's own last-gen Flash.
  • Stay on Claude Fable 5 for the deepest coding and reasoning work. Nothing in today's launch competes at the top of the index, and Google is not claiming otherwise.
  • Stay on GPT-5.5 if you need the strongest broadly available non-Claude reasoning. Same logic.
  • Gemini 3.5 Flash Cyber is not a choice you can make. Move on.

The honest summary: Google shipped a same-IQ model on purpose, and for the agentic workloads it targets, that was probably the right call. Faster and cheaper at equal quality is a real upgrade, just not the kind that makes headlines. The headline-making model is apparently called Gemini 4, and it is in the oven.


Sources