Gemini 3.7 Flash Is Smarter and Half Price. Read the Fine Print on Both.
Google shipped Gemini 3.7 Flash on August 13, just 23 days after 3.6, and this time the independent numbers actually moved: 56 on the Artificial Analysis index, up 4 points. The half price is real too, until December 31. We verified the pricing page, the model card, and the benchmark table, and found the asterisks Google does not lead with.
By T.W. Ghost
Three weeks ago, Google shipped Gemini 3.6 Flash with the same intelligence score as its predecessor and made efficiency the whole pitch. We called it an unusual thing for an AI lab to admit. Twenty-three days later, Gemini 3.7 Flash is the other kind of launch: the independent numbers moved.
It landed August 13, the closing act of a very loud week for AI agents. Two days earlier, xAI had shipped Grok Bot and then Grok 4.6, a story with its own fine print that we cover separately. Google's answer is not a worker, it is a cheaper, smarter engine. Here is what survived verification.
This Time It Actually Got Smarter
Artificial Analysis, which runs its own benchmarks rather than reprinting vendor decks, scores 3.7 Flash at 56 on its Intelligence Index at high thinking, against 52 for 3.6 Flash. A four-point independently measured gain in three weeks, at a measured 340.1 tokens per second output speed and a blended cost of $0.58 per million tokens, roughly half of 3.6 Flash's $1.16 at its original price. Google's own model card is unusually direct about the lineage: 3.7 Flash "is based on Gemini 3.6 Flash."
| Gemini 3.7 Flash | |
|---|---|
| Announced | August 13, 2026, GA (no preview suffix) |
| Model ID | gemini-3.7-flash |
| API price (intro) | $0.75 in / $3.75 out per 1M, through Dec 31, 2026 |
| API price (standard) | $1.50 in / $7.50 out per 1M, from Jan 1, 2027 |
| Context | 1,048,576 tokens in / 65,536 out |
| Modalities | Text, image, video, audio, PDF in; text out |
| Knowledge cutoff | March 2026 (some domains only to Jan 2025, per the model card) |
| Thinking levels | Low, medium, high. "Minimal" now returns an API error |
| AA Intelligence Index | 56 (high thinking), vs 52 for 3.6 Flash |
| AA output speed | 340.1 tok/s measured |
One small receipt while we are here: several outlets, including two that reviewed the model at length, printed the model ID as gemini-3-7-flash with hyphens. Google's API docs say gemini-3.7-flash, dots, matching every prior Gemini model code. If your API call 404s, that is why.
The Half-Price Asterisk
The headline everywhere is "50% price cut." The pricing page tells a more precise story, and it is worth two minutes of your time if you are budgeting an agent pipeline.
The intro price is a sale with an end date. $0.75 in / $3.75 out is what Google's pricing page calls an introductory rate, in effect through December 31, 2026. On January 1, 2027 it doubles to $1.50/$7.50, which is exactly what 3.6 Flash cost at launch. Caching follows the same curve ($0.075 rising to $0.15 per 1M, cache storage $0.50 per 1M per hour rising to $1.00), and so does Batch ($0.375/$1.875 rising to $0.75/$3.75).
The old model got the same price the same day. Google simultaneously repriced Gemini 3.6 Flash to the identical $0.75/$3.75. So the "50% cut" is real only against 3.6 Flash's original launch price. Today the two models bill identically, and the honest framing is: the smarter model costs the same as the old one, and both are on sale until New Year's Eve.
A rhyme we cannot resist: in our 3.6 Flash coverage we noted the Batch and Flex tiers were half price at $0.75/$3.75. That discount tier is now the standard rate. The sale price was sitting in the price list all along.
Thinking tokens bill as output. Google's pricing page bills "output (including thinking)" at the output rate, and the API returns only a summary of the thinking. You pay full freight for tokens you mostly do not see. The "minimal" thinking level from earlier Flash models is gone entirely and returns an API error, which several reviewers flagged as a real regression for high-volume classification workloads that wanted zero reasoning overhead.
For scale: the intro price undercuts Claude Haiku 4.5 ($1.00/$5.00) on both sides. But GPT-5.6 Luna sits at $0.20/$1.20 after OpenAI's 80% July cut, and DeepSeek V4-Flash is at $0.14/$0.28. Cheap is a crowded neighborhood.
The Benchmarks, Sorted by Who Reported Them
House rules: Artificial Analysis runs its own harness, so its numbers get stated as fact with citation. Everything else in this section is Google's self-reported launch numbers, printed from Google's model card, and labeled accordingly.
Google reports, versus 3.6 Flash:
| Benchmark | 3.7 Flash | 3.6 Flash |
|---|---|---|
| FrontierCode 1.1 Main | 43.6% | 34.4% |
| DeepSWE v1.1 | 65.3% | 48.6%* |
| WebDev Arena Elo | 1588 | 1538 |
| Terminal-bench 2.1 | 85.8% | 78.0% |
| Terminal-bench 3.0 | 14.9% | 5.4% |
| OSWorld-2.0 (computer use) | 47.9% | 33.8% |
| GDP.pdf (documents) | 34.0% | 22.0% |
| AutomationBench | 30.4% | 17.0% |
| GDPval-AA v2 | 1525 Elo | 1422 Elo |
*Google's blog post says the 3.6 baseline was 49.0%. Google's own model card says 48.6%. We print the model card and note the discrepancy, because that is the whole job.
Three honest observations Google's marketing does not lead with:
- The model card shows a regression. CharXiv chart reasoning went from 85.2% to 84.5%. Tiny, but it is there, and the coverage that reprinted the launch table skipped it.
- On Google's own comparison table, Google does not sweep. The model card compares 15 benchmark rows against rivals; 3.7 Flash wins 8. GPT-5.6 Terra still leads the agentic headliners: DeepSWE v1.1 (69.6%), Terminal-bench 2.1 (87.4%), Terminal-bench 3.0 (20.8%), and OSWorld-2.0 (50.2%).
- AutomationBench at 30.4% means the model still fails about seven in ten multi-step automation tasks. An 80% improvement on a benchmark can coexist with "do not ship this unattended."
And the caveat that carries over from 3.6 Flash unchanged: Artificial Analysis measures time to first token at 9.83 seconds at high thinking. Blistering throughput once it starts, terrible for a live chat box. This model is built for agents and pipelines, which is exactly what Google says it is for.
For the wider standings: Grok 4.6, released the day before, scores 61 on the same index, and the frontier is a Claude story (Opus 5 at 63, Fable 5 at 62). But none of those are in this price class. Nothing at 3.7 Flash's price scores higher than 3.7 Flash right now, until you drop to the true budget tier where Luna and DeepSeek live.
Where You Can Run It
The Gemini API in AI Studio, Android Studio, Google Antigravity, the Gemini Enterprise Agent Platform, and Gemini Spark, Google's always-on personal agent for AI Pro and Ultra subscribers in 160+ countries. Spark notably excludes the EEA, the UK, Switzerland, and Nigeria, and has no free-tier route. Computer use is in preview, and Search and Maps grounding work at launch.
One roadmap note, because we track our own record: Gemini 3.5 Pro, announced at I/O in May, has now missed three reported release windows and remains unshipped. The newest released Pro model is still 3.1 Pro at $2/$12. Google keeps shipping Flash models into the gap, and at these prices, that may be the strategy.
Verified vs Unconfirmed: The Scorecard
| Claim | Verdict |
|---|---|
| Intro price $0.75/$3.75 through Dec 31, 2026, then $1.50/$7.50 | Verified (Google's pricing page) |
| Gemini 3.6 Flash repriced to identical rates; "50% cut" is vs its original price | Verified (Google's pricing page) |
| AA Intelligence Index 56 vs 52, 340.1 tok/s, $0.58 blended, 9.83s first token | Verified (Artificial Analysis) |
| Google's benchmark table, incl. the CharXiv regression and Terra's agentic leads | Verified (DeepMind model card) |
Model ID is gemini-3.7-flash with dots, not hyphens | Verified (Google API docs; two review sites printed it wrong) |
| "Fastest of 186 models tracked" speed ranking | Unverified (one outlet's characterization; AA's own page says "well above average") |
| Gemini 3.5 Pro still unshipped after its May announcement | Verified (Bloomberg, Axios, 9to5Google) |
Who Should Use It
Use Gemini 3.7 Flash if you run coding and automation agents at volume, where cost per task beats peak intelligence and nobody is watching a chat window. It is the best price-to-capability ratio Google has ever shipped, until December 31. Budget against the standard price, not the sale price, or January will be an unpleasant sprint.
Skip it for latency-sensitive chat (10 seconds to first token), for high-volume zero-reasoning classification (minimal thinking is gone), and for the hardest agentic workloads, where GPT-5.6 Terra leads on Google's own table and Grok 4.6 posts the best measured cost per task at a higher intelligence tier.
The frontier is elsewhere: Claude Opus 5 (63) and Claude Fable 5 (62) for the hardest reasoning and deepest coding, with Kimi K3 holding the open-weights flag at 60. Our model comparison has the current standings side by side.
Sources
- Google: Introducing Gemini 3.7 Flash
- Google Gemini API pricing page
- Google Gemini API model docs: gemini-3.7-flash
- DeepMind model card: Gemini 3.7 Flash
- Artificial Analysis: Gemini 3.7 Flash
- 9to5Google: Gemini 3.7 Flash launch
*Not sure whether the cheap agent engine fits how you actually work, or whether you need a different lab entirely? Take the free 2-minute quiz and get matched. Then read the other half of this week's agent story: Grok Bot and Grok 4.6.*