AERIOXFLUX
← Frontier Labs
Frontier Labs · google deepmind

Gemini 4 Argon Ships to Cyber Defenders Before Developers

Google DeepMind's first Gemini 4 model can write a million tokens in one answer and hallucinates far less than its rivals, but the first people allowed to use it are security teams, not paying customers.

Flux Desk·2026-10-01·5 min read

Google DeepMind's first Gemini 4 model arrived on September 30 with a spec that would have sounded absurd a year ago: a single response can run to one million tokens. Gemini 4 Argon, announced on Google's blog by Koray Kavukcuoglu, SVP of Google DeepMind and Google's Chief AI Architect, raises the output ceiling from 64,000 tokens and pitches itself as a model for long-horizon work in software engineering, legal and financial analysis, and cyber defense.

The more interesting decision is who gets it first. Argon is not in the API. It is rolling out to what Google calls the Fairwind Program, a cohort of trusted cyber defenders and government users, and only later to paid API customers and Google AI Ultra subscribers "as soon as possible." Developers who want to test it this week cannot.

A million tokens, with an asterisk

The 1M-token output limit is the headline feature, and it changes what a single call can do: a full codebase migration, a long legal memo with exhibits, or hours of agentic reasoning without stitching together separate requests. Google says it is already using Argon internally on exactly that kind of job. According to the company's post, the model has freed more than 300 TiB of memory across Google's data centers, with an estimated eventual saving of 500 TiB to 1 PiB, and is working through more than 800,000 lines of code in a migration of the Fuchsia Zircon kernel.

The number deserves a footnote. Latent Space's AINews reported that independent evaluator Vals measured a maximum of 262,000 output tokens in practice, with the full million reached through a feature Google calls Long Decode Continuation. That is still four times the old limit. It is not quite the same as a model that emits a million tokens in one uninterrupted pass.

The benchmark picture is mixed, and honest about it

On Google's own chart, Argon posts 77.9% on DeepSWE v1.1, ahead of Claude Opus 5.5 at 74.2% and OpenAI's GPT-6 Astra at 74.1%, according to DataCamp's breakdown of the published numbers. It leads AutomationBench at 51.3%, ties Astra at 68% on the CWE-bench v1 vulnerability-remediation test, and scores 91.7% on LVBench for long video. The domain results are wider still: 19.6% on Harvey's Legal Agent Benchmark against 6.7% for Claude Fable 5.1, and 65.4% on Vals Finance Agent v2.

Google also published the losses. On FrontierSWE v2, Astra scores 65.5% and Opus 5.5 62.3%, while Argon trails at 55.0%. On Terminal-bench 4.0, Opus 5.5 leads at 66.4%, with Astra at 58.2% and Argon at 57.4%. Latent Space tallied Argon as state of the art on 13 of 19 credible benchmarks, which is a strong showing and not a sweep. It also placed Argon first on the Text Arena leaderboard at 1525.

Independent scoring puts the frontier in a dead heat. On the Artificial Analysis Intelligence Index, per Yahoo Tech's report, Argon scores 53, level with GPT-6 Astra, one point above GPT-6.1 Sol at 52 and well clear of Grok 4.7 at 46. Gemini 3.1 Pro Preview sat at 30 on the same index, which shows how large the generational jump is inside Google.

The real differentiator is what it refuses to make up

The number most likely to matter to enterprise buyers is not a capability score. On Artificial Analysis's AA-Omniscience test, Argon hallucinated at a 15% rate, against 51% for Astra, according to Yahoo Tech's summary of the results. The trade is visible in the same data: Argon answered 50% of questions correctly, 13 points behind Astra's 63%.

Put plainly, Argon knows less on that test but is far more willing to say so. For legal and finance work, which Google names as core targets, a model that abstains is worth more than one that answers confidently and wrong. Google appears to be betting that buyers in regulated industries will choose calibration over raw recall, and its decision to lead with Harvey's legal benchmark and a finance agent eval reinforces that.

Cost cuts the other way. Argon averaged about 62,000 output tokens per task on the Artificial Analysis index, versus 27,000 for Astra. At introductory prices that still makes Argon cheaper to run the full index, $1.99 against Astra's $3.26, but at standard pricing the figure rises to $3.98.

Pricing with an expiry date

Google is launching Argon at $2 per million input tokens and $10 per million output tokens, with cached input discounted 95%. That is a 50% introductory discount. The standard rate is $4 input and $20 output, and Google has not said when the introductory window closes. For a model built to produce very long outputs, the output price is the one that matters, and it doubles on a date nobody has announced.

There is also no public model ID yet. DataCamp noted that as of launch day Argon was not listed on OpenRouter, Vertex AI or GitHub Copilot, which reinforces how staged this release is.

Why defenders go first

Google's stated reasoning is capability risk. The blog post says that for "trusted defenders and our own internal teams at Google, we'll be releasing Argon without cyber guardrails." The public version will refuse harmful requests while trying to preserve legitimate dual-use security research. Google cites Wiz, through its Scan for Good initiative, as an early user, and says Argon found a critical healthcare vulnerability that earlier frontier models missed.

The approach mirrors what other labs have done with their most capable models this year: restrict the first release to vetted security users, collect real-world evidence, then widen access. Google also used the post to make a pointed request of competitors, writing that it "strongly encourage[s] the rest of the industry to preserve reasoning transparency" as capabilities rise. That is a statement about chain-of-thought monitoring, which Google says it applies to Argon's reasoning and actions.

The open question is how the model behaves outside the evaluation harness. A defender-first rollout gives Google time and a controlled user base to answer it. It also means that for now, the most capable Gemini ever built is something most developers can read about but cannot call.

#gemini 4#argon#google deepmind#hallucination#frontier models

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.