AERIOXFLUX
Frontier Labs
Frontier Labs · benchmarks safety

Hassabis Wants Your Model 30 Days Before You Ship It

DeepMind's CEO took his proposal for an independent frontier-AI standards body directly to the Treasury Secretary — and the mechanism he is pitching has teeth.

Flux Desk·2026-08-14·5 min read

Demis Hassabis has spent much of 2026 arguing publicly for an independent body to set safety standards for frontier AI. This week the argument moved out of op-eds and conference stages: reporting on August 13, 2026 describes Hassabis personally pitching the concept to Treasury Secretary Scott Bessent and OSTP Director Michael Kratsios.

He wants it operational before the end of 2026.

The mechanism

Stripped of the diplomacy, the proposal has a specific and demanding shape.

Frontier labs would submit their most capable models to an independently governed body up to 30 days before public release. The body would test those models against common standards, and would have the authority to intervene if a system displayed dangerous capabilities. Funding would come from industry; governance would not.

Hassabis has described the eventual international version by analogy to the International Atomic Energy Agency — the UN-linked organization that combines technical inspection with a safeguards regime. In earlier framings he has also invoked FINRA, the industry-funded self-regulator for US securities firms.

Those two analogies point in noticeably different directions, and the difference is the substance of the whole debate.

IAEA or FINRA

FINRA is a self-regulatory organization. The industry funds it, the industry largely staffs it, and its authority derives from the fact that regulators would otherwise do the job less sympathetically. It is fast, technically competent, and structurally disinclined to take actions that seriously damage its members.

The IAEA is an intergovernmental body with treaty authority and inspection rights that member states have formally ceded. It is slower and more political, and it can say no in ways a trade association cannot.

The proposal on the table has FINRA's funding model and IAEA's ambitions. Industry-funded, independently governed, with intervention authority over product launches.

That combination is not impossible. It is, however, the exact structure whose credibility depends entirely on governance details that have not been published — who appoints the board, who can remove them, what the appeal process is when a lab is told to delay a launch, and what happens when the body's largest funders are also its subjects.

Why thirty days is the interesting number

Pre-release testing windows have been discussed for two years. Thirty days is a specific choice, and it is short.

It is too short to do genuinely novel safety research on a new model. Discovering an unanticipated capability class in a frontier system — the kind of finding that takes red teams months — does not fit in a month.

It is long enough to run an established battery of evaluations against known dangerous-capability thresholds: bioweapon uplift, offensive cyber capability, autonomous replication, deception under evaluation. That is the actual design target. The body is not meant to discover new categories of risk. It is meant to check, consistently and independently, whether a given model crosses lines the field has already agreed matter.

Thirty days is also a number labs might accept. A ninety-day window would be a competitive weapon — three months is long enough for a rival to ship first, and no lab in a race this tight will hand that to a third party. A month is short enough to survive a product calendar.

Which tells you the proposal was designed to be adopted rather than to be maximally protective. Whether that is pragmatism or capture depends on where you stand.

The strategic reading

It is worth being clear-eyed about incentives. Hassabis leads a frontier lab. A pre-release testing regime with intervention authority imposes real costs on every frontier lab — including his own.

It imposes them asymmetrically.

An organization with an established safety team, a published risk framework, and existing evaluation infrastructure absorbs a mandatory testing window as process overhead. It already does most of this internally. An organization moving fast with a thin safety function absorbs it as a genuine constraint on velocity.

Google DeepMind is firmly in the first category. So is Anthropic. So is OpenAI. The labs that would feel this most are the fast followers, the well-funded newcomers, and the open-weight developers whose release model does not have a clean "pre-release" moment at all.

That is not an argument the proposal is bad faith. Hassabis has been consistent on frontier risk for a decade, well before it was commercially convenient. It is an argument that a rule can be sincerely motivated and still advantage the people proposing it, and that both facts should be held at once.

The complication in Washington

The reporting includes a detail that complicates the picture considerably: while Hassabis was lobbying Treasury for an industry-funded independent body, Treasury was reportedly developing its own approach.

That is the fault line this proposal has to cross. Every industry-led standards body in history has been a negotiation over whether government delegates authority or exercises it. The industry offers speed, technical competence, and voluntary compliance. Government retains legitimacy, enforcement power, and the ability to bind non-participants.

The awkward truth is that an industry-funded body cannot compel a lab that refuses to join, and the labs most likely to refuse are exactly the ones the regime exists to constrain. A voluntary regime that captures the four most safety-conscious labs and misses everyone else has produced coordination among the careful, which was not the problem.

What to watch

Hassabis wants this standing before the end of the year. Four months is an extremely aggressive timeline for an institution of this kind, and it suggests he believes the window for industry-led structure is closing.

Three signals will tell you whether it happens.

Whether any lab other than Google DeepMind publicly commits to the 30-day submission. One lab proposing a rule is a position; three labs accepting it is an institution.

Whether the governance documents get published before the body exists. Independence is a claim until the appointment and removal mechanisms are on paper.

And whether Treasury's own effort proceeds in parallel. If it does, the question stops being what the right regime looks like and becomes who gets to build it — which is a much older kind of fight, and one that industry usually loses slowly.

#hassabis#deepmind#ai-safety#regulation#standards-body

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.