Google DeepMind
Gemini 3.7 Flash
the workhorse tier, on an introductory price

Key facts
- 13 Aug 20263 weeks after 3.6 Flash
- Released
- 1M / 64ktokens in / out
- Context
- $0.75 / $3.75per M, until 31 Dec 2026
- Price
- $1.50 / $7.50standard rate
- From 1 Jan 2027
- 5617th of 188 tracked
- AA Index
- 340 tok/sfastest of 188 tracked
- Speed
A fine-tune of 3.6 Flash shipped three weeks after it, clearly better at coding and agent work on Google's own table, and the fastest model Artificial Analysis has measured. The catch sits in the price: the launch rate is introductory, and doubles on 1 January 2027.
What it is
Gemini 3.7 Flash is Google DeepMind’s workhorse model, announced on 13 August 2026, three weeks after Gemini 3.6 Flash. Flash is the tier for high-volume everyday work, and the announcement, bylined by product management director Tulsee Doshi, calls it “our most intelligent workhorse model yet for coding and agents”.
It is an iteration rather than a new foundation, and Google says so: the model card states that “Gemini 3.7 Flash is based on Gemini 3.6 Flash”, with “algorithmic improvements to its core reasoning foundation” and “customizable thinking configurations to control the mix of quality, cost and latency”. The release cadence is the striking part. This is the third new Flash model in under a month and a half, after 3.6 Flash and 3.5 Flash-Lite arrived together on 21 July.
What Google’s own table shows
Google published its comparison against 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra and Muse Spark 1.2 across twenty rows, and it rewards a close read, because it is not a clean sweep. A benchmark is a fixed set of tasks producing a comparable score, and the publisher of a launch table has a stake in it. Ten representative rows below, with the leader named:
| Benchmark | 3.7 Flash | 3.6 Flash | Leader |
|---|---|---|---|
| AA Intelligence Index | 56 | 52 | 57 GPT-5.6 Terra / Muse Spark |
| FrontierCode 1.1 Main | 43.6% | 34.4% | 3.7 Flash |
| DeepSWE v1.1 | 65.3% | 48.6% | 69.6% GPT-5.6 Terra |
| Code Arena (Elo) | 1588 | 1538 | 3.7 Flash |
| Terminal-bench 3.0 | 14.9% | 5.4% | 20.8% GPT-5.6 Terra |
| GDPVal-AA v2 (Elo) | 1525 | 1422 | 1628 Muse Spark 1.2 |
| GDP.pdf | 34.0% | 22.0% | 3.7 Flash |
| Agent’s Last Exam | 26.3% | 24.2% | 33.3% Claude Sonnet 5 |
| GDM-MRCR v2 (128k) | 97.0% | 91.8% | 3.7 Flash |
| AutomationBench* | 30.4% | 17.0% | 3.7 Flash |
* Flagged by Google itself as a private, non-reproducible test set.
Counted up, 3.7 Flash takes the best score on nine of the twenty rows, GPT-5.6 Terra on seven, Claude Sonnet 5 on two, Muse Spark 1.2 on the knowledge-work Elo plus a share of the index, and 3.6 Flash keeps one row against its own successor, tool-assisted chart reasoning. Against its own predecessor the step is large wherever code or agents are involved: DeepSWE jumps from 48.6 to 65.3 per cent, FrontierCode from 34.4 to 43.6, Terminal-bench 3.0 nearly triples. Against the field, the workhorse claims hold best on document work, long context and web development, while GPT-5.6 Terra leads most of the agentic rows and Claude Sonnet 5 takes Agent’s Last Exam by seven points. Google printed all of this itself.
The independent reading agrees on the headline number: Artificial Analysis scores 3.7 Flash at 56 on its Intelligence Index, seventeenth of the 188 models it tracks. Where it stands first of all 188 is speed, at 340.1 output tokens per second.
The price has a clock on it
The launch price is $0.75 per million input tokens and $3.75 per million output, half of what 3.6 Flash launched at three weeks earlier. The footnote is doing heavy lifting: this is an introductory rate that runs to 31 December 2026, after which the standard rate of $1.50 and $7.50 applies. Context caching, batch and the rest follow the same pattern, half price now, doubling in January.
There is a free tier on the API, with a difference worth knowing: on the free tier Google answers “yes” to whether prompts are “used to improve our products”, and “no” on the paid one.
The small print
The knowledge cutoff is stated as March 2026, with an unusual caveat quoted directly from the model card: in some domains “the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family)”. Inputs cover text, images, audio, video and PDF; output is text only, up to 64,000 tokens, with a one-million-token input window.
On safety, the model card reports no tracked or critical capability level reached under Google’s Frontier Safety Framework, while noting that CBRN and cybersecurity evaluations both “reach the alert threshold” for Level 1. The full safety report was not published at launch; the card says it “will be published shortly”.
It is available in the Gemini app through Gemini Spark for AI Pro and Ultra subscribers in over 160 countries, in Google AI Studio, on the Gemini API, and inside Antigravity and Android Studio. Whether the free consumer Gemini app has moved to 3.7 Flash is something Google has not said.
For the tier it replaces at the top of the Flash line, see Gemini 3.6 Flash; for the family picture, the Gemini 3.5 family.
More in Large Language Models
All LLMs →- OpenAIGPT-5.6the flagship since 9 July 2026
- AnthropicClaude Fable 5the flagship topping most July 2026 rankings
- AnthropicClaude Opus 5frontier work with a dial on the bill
- AnthropicClaude Mythos 5restricted twin of Fable 5
- AnthropicClaude Sonnet 5the speed and intelligence balance
- Google DeepMindGemini 3.5 familythe generation behind Gemini today