Google DeepMind

Gemini 3.7 Flash

the workhorse tier, on an introductory price

Released 13 August 20265 min readLarge Language ModelsLast updated:

Editorial collage: the Gemini spark mark beside a workbench spec card reading 1 million token context, 64,000 out, knowledge cutoff March 2026, a price tag reading $0.75 in and $3.75 out per million until 31 December 2026 then $1.50 and $7.50, and a speedometer needle at 340 tokens per second over a halftone data-centre hall

Key facts

13 Aug 20263 weeks after 3.6 Flash
Released
1M / 64ktokens in / out
Context
$0.75 / $3.75per M, until 31 Dec 2026
Price
$1.50 / $7.50standard rate
From 1 Jan 2027
5617th of 188 tracked
AA Index
340 tok/sfastest of 188 tracked
Speed

A fine-tune of 3.6 Flash shipped three weeks after it, clearly better at coding and agent work on Google's own table, and the fastest model Artificial Analysis has measured. The catch sits in the price: the launch rate is introductory, and doubles on 1 January 2027.

What it is

Gemini 3.7 Flash is Google DeepMind’s workhorse model, announced on 13 August 2026, three weeks after Gemini 3.6 Flash. Flash is the tier for high-volume everyday work, and the announcement, bylined by product management director Tulsee Doshi, calls it “our most intelligent workhorse model yet for coding and agents”.

It is an iteration rather than a new foundation, and Google says so: the model card states that “Gemini 3.7 Flash is based on Gemini 3.6 Flash”, with “algorithmic improvements to its core reasoning foundation” and “customizable thinking configurations to control the mix of quality, cost and latency”. The release cadence is the striking part. This is the third new Flash model in under a month and a half, after 3.6 Flash and 3.5 Flash-Lite arrived together on 21 July.

Screenshot of Google's announcement headed Introducing Gemini 3.7 Flash, subtitled Our most intelligent workhorse model yet for coding and agents, bylined Tulsee Doshi, dated 13 August 2026
Google's announcement, 13 August 2026. Screenshot of blog.google, 14 August 2026.

What Google’s own table shows

Google published its comparison against 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra and Muse Spark 1.2 across twenty rows, and it rewards a close read, because it is not a clean sweep. A benchmark is a fixed set of tasks producing a comparable score, and the publisher of a launch table has a stake in it. Ten representative rows below, with the leader named:

Benchmark 3.7 Flash 3.6 Flash Leader
AA Intelligence Index 56 52 57 GPT-5.6 Terra / Muse Spark
FrontierCode 1.1 Main 43.6% 34.4% 3.7 Flash
DeepSWE v1.1 65.3% 48.6% 69.6% GPT-5.6 Terra
Code Arena (Elo) 1588 1538 3.7 Flash
Terminal-bench 3.0 14.9% 5.4% 20.8% GPT-5.6 Terra
GDPVal-AA v2 (Elo) 1525 1422 1628 Muse Spark 1.2
GDP.pdf 34.0% 22.0% 3.7 Flash
Agent’s Last Exam 26.3% 24.2% 33.3% Claude Sonnet 5
GDM-MRCR v2 (128k) 97.0% 91.8% 3.7 Flash
AutomationBench* 30.4% 17.0% 3.7 Flash

* Flagged by Google itself as a private, non-reproducible test set.

Counted up, 3.7 Flash takes the best score on nine of the twenty rows, GPT-5.6 Terra on seven, Claude Sonnet 5 on two, Muse Spark 1.2 on the knowledge-work Elo plus a share of the index, and 3.6 Flash keeps one row against its own successor, tool-assisted chart reasoning. Against its own predecessor the step is large wherever code or agents are involved: DeepSWE jumps from 48.6 to 65.3 per cent, FrontierCode from 34.4 to 43.6, Terminal-bench 3.0 nearly triples. Against the field, the workhorse claims hold best on document work, long context and web development, while GPT-5.6 Terra leads most of the agentic rows and Claude Sonnet 5 takes Agent’s Last Exam by seven points. Google printed all of this itself.

Screenshot of Google's own comparison table for Gemini 3.7 Flash against Gemini 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra and Muse Spark 1.2, including price rows and the Artificial Analysis Intelligence Index row reading 56, 52, 55, 57, 57
The comparison as Google presents it, price rows included. Screenshot of deepmind.google, 14 August 2026.

The independent reading agrees on the headline number: Artificial Analysis scores 3.7 Flash at 56 on its Intelligence Index, seventeenth of the 188 models it tracks. Where it stands first of all 188 is speed, at 340.1 output tokens per second.

The price has a clock on it

The launch price is $0.75 per million input tokens and $3.75 per million output, half of what 3.6 Flash launched at three weeks earlier. The footnote is doing heavy lifting: this is an introductory rate that runs to 31 December 2026, after which the standard rate of $1.50 and $7.50 applies. Context caching, batch and the rest follow the same pattern, half price now, doubling in January.

The introductory clock. Budgeting on the launch price means budgeting for a doubling in January.

There is a free tier on the API, with a difference worth knowing: on the free tier Google answers “yes” to whether prompts are “used to improve our products”, and “no” on the paid one.

The small print

The knowledge cutoff is stated as March 2026, with an unusual caveat quoted directly from the model card: in some domains “the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family)”. Inputs cover text, images, audio, video and PDF; output is text only, up to 64,000 tokens, with a one-million-token input window.

On safety, the model card reports no tracked or critical capability level reached under Google’s Frontier Safety Framework, while noting that CBRN and cybersecurity evaluations both “reach the alert threshold” for Level 1. The full safety report was not published at launch; the card says it “will be published shortly”.

It is available in the Gemini app through Gemini Spark for AI Pro and Ultra subscribers in over 160 countries, in Google AI Studio, on the Gemini API, and inside Antigravity and Android Studio. Whether the free consumer Gemini app has moved to 3.7 Flash is something Google has not said.

For the tier it replaces at the top of the Flash line, see Gemini 3.6 Flash; for the family picture, the Gemini 3.5 family.