AERIOXFLUX
Frontier Labs
Frontier Labs · google deepmind

Gemini 3.7 Flash Is Half Price Until Google Says Otherwise

Three weeks after 3.6, Google shipped a Flash model with double-digit benchmark jumps at $0.75 in and $3.75 out — with list prices contractually doubling on January 1.

Flux Desk·2026-08-15·5 min read

Google released Gemini 3.7 Flash on August 13, 2026, three weeks after Gemini 3.6 Flash. Through December 31, 2026, it costs $0.75 per million input tokens and $3.75 per million output tokens. On January 1, 2027, list prices double to $1.50 and $7.50.

The introductory rate halves what Flash-class models have cost. The expiry date is printed on the tin.

The benchmark movement is real

Three weeks between releases usually means a checkpoint refresh. The numbers say otherwise.

On FrontierCode 1.1 Main, 3.7 Flash scores 43.6%, up from 34.4% for 3.6. On DeepSWE v1.1 it scores 65.3%, up from 49.0%. On Code Arena it takes an Elo of 1588, against 1538 for 3.6 Flash, 1541 for Claude Sonnet 5, and 1523 for GPT-5.6 Terra.

A sixteen-point absolute gain on DeepSWE in twenty-one days is not a checkpoint. That is a training or post-training change of real substance, shipped at a cadence the rest of the industry is not matching on its workhorse tier.

The context window stays at 1 million tokens.

The losses are the informative part

Google published the benchmarks it does not win, which makes the ones it does win more credible.

On Terminal-bench 2.1, 3.7 Flash lands at 85.8%, behind GPT-5.6 Terra at 87.4%. On the multimodal Agent's Last Exam, Claude Sonnet 5 leads at 33.3% against Flash's 26.3%.

Read together, the picture is specific rather than general. Gemini 3.7 Flash is at or near the frontier on code generation and repair while trailing on terminal-driven agentic execution and multimodal agent reasoning. Those are different jobs. A model that writes and fixes code exceptionally well but is a step behind on long-horizon tool use is a superb component inside a harness someone else designs, and a weaker choice as the autonomous driver of one.

That is a coherent product position for a model explicitly aimed at coding, agentic workflows, and document processing in enterprise settings.

Half price with a fuse

The pricing structure deserves more attention than the benchmarks.

Google is not cutting prices. It is running a promotion with a stated end date, after which the rate doubles. Every enterprise that builds on 3.7 Flash between now and December is building a cost model that breaks on New Year's Day unless something else changes by then.

There are two readings, and both are probably partially true.

The customer-acquisition reading is that Google is buying the fourth quarter — the period when enterprises run pilots, finalize next year's budgets, and pick default models. Winning that window at half price and then repricing into signed commitments is a well-established playbook, and it works because switching costs accrue faster than procurement cycles.

The capacity reading is that Google knows what its serving costs will look like in six months and has priced accordingly, betting that efficiency gains or a successor model will absorb the difference before customers feel it. If 3.8 Flash arrives in November at the promotional rate, the January step-up quietly never applies to anyone paying attention.

Either way, the honest characterization is that $0.75/$3.75 is a marketing price and $1.50/$7.50 is the product's price. Model your 2027 spend on the second number.

The cadence is the strategic weapon

Three weeks between Flash releases, against an industry norm measured in months, changes the competitive dynamic more than any single benchmark.

A competitor that ships a model beating Gemini 3.7 Flash in October is not beating Google's current model by the time procurement closes — it is beating a model two revisions old. Sustained fast cadence converts a benchmark race into an attrition contest, and attrition contests are won by whoever owns their own silicon, their own data centers, and their own distribution.

There is a real cost to it, though, and enterprises are the ones who pay it. Every model swap means re-running evaluations, re-tuning prompts, re-validating outputs against compliance requirements, and re-certifying anything that touches a regulated workflow. A three-week cadence is a gift to a startup iterating on a demo and a tax on a bank with a model risk committee.

The read

Gemini 3.7 Flash is a genuine capability jump on coding benchmarks, at a price that is temporarily aggressive and permanently reasonable, from a vendor shipping faster than anyone else in the tier.

The strategic move underneath it is price-anchoring the workhorse category. Frontier models get the headlines; Flash-class models get the volume, because the overwhelming majority of enterprise tokens go to tasks that do not need a frontier model. Whoever sets the reference price for that tier sets the economics of the whole agentic software category.

Google just set it at $0.75, with a note saying the real number is $1.50 and everyone has until January to get comfortable.

The competitors now have to answer the promotional price, not the list price. That is the entire point.

#google#gemini#model-pricing#coding-agents#benchmarks

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.