AERIOXFLUX
Frontier Labs
Frontier Labs · benchmarks safety

FrontierMath Tier 4 Is Finished, and It Took 14 Months

Every Tier 4 problem has now been solved, with GPT-6 Astra taking the last one standing. The benchmark opened at 5% in July 2025. It closed at 98%.

Flux Desk·2026-09-11·5 min read

Epoch AI announced that every problem in FrontierMath Tier 4 has now been solved by AI. GPT-6 Astra took the last one standing — a problem authored by Jay Pantone.

Tier 4 launched on July 11, 2025. The top score at launch was 5%. Less than 14 months later the top score is 98% and Epoch considers the benchmark saturated.

FrontierMath Tier 4 was built to be the hard one. Tiers 1 through 3 spanned undergraduate to graduate difficulty; Tier 4 was research-level problems commissioned from working mathematicians, designed so that a specialist in an adjacent field would struggle. The explicit intent was a benchmark with years of headroom.

It had fourteen months.

The detail that matters most

Epoch noted something specific about the final problem: mathematicians frequently observed that AI found unintended shortcuts when solving their Tier 4 problems. Not so for this last one.

That single qualification carries more weight than the 98%.

A large share of benchmark progress in mathematical reasoning has been contaminated by exactly this failure mode. A problem is designed to require a deep insight; the model finds a computational route the author did not anticipate, or exploits structure in how the answer is checked, and produces the correct value without the reasoning the problem was meant to test. The score goes up. The capability being measured does not.

Pantone's problem apparently resisted that. Which means the last data point on the curve is the cleanest one — and the one that most supports reading the saturation as genuine capability rather than accumulated benchmark gaming.

What 5% to 98% in 14 months actually measures

Three things happened simultaneously, and they are worth separating because they have different implications.

Reasoning-model scaling. The period from mid-2025 to late 2026 is the era of inference-time compute becoming a first-class scaling axis. Tier 4 problems reward long, structured, verifiable chains of work — precisely the shape of task that benefits most from spending more tokens thinking. The curve is steep partly because the benchmark sat directly in the path of the dominant capability trend.

Agent orchestration. The Navier-Stokes result OpenAI published days earlier used roughly 10,000 concurrent agents, 2.7 million messages, and 130 billion output tokens across 88 hours, followed by 17 hours of formal verification in Lean. Research-level mathematics turned out to be extremely amenable to parallel search with verification, and that is an orchestration achievement as much as a model one.

Formal verification as a grading mechanism. When output can be machine-checked in Lean, reinforcement learning gets a clean reward signal. Mathematics is one of the few domains where correctness is decidable, which makes it one of the few domains where you can grind capability upward without a human in the loop. The pace on math benchmarks was always going to outrun the pace on domains where grading requires judgment.

Why this is a problem for everyone measuring progress

Tier 4 was among the last benchmarks anyone trusted to have durable headroom. Its saturation leaves the evaluation ecosystem in a bad position.

Look at what happened to Terminal-Bench in the same week. Cognition's SWE-2 scored 92.8% on Terminal-Bench 2.1 and 27.3% on Terminal-Bench 4 — a 65-point gap on the same benchmark family, because 2.1 has been optimized against by everyone for long enough to be meaningless while 4 is new. CursorBench 4.0 raised difficulty and scores fell across every model. Epoch has already shipped a Tier 4 v2.

The pattern is now the norm: a benchmark is useful for roughly a year, gets saturated, gets replaced, and every cross-version comparison in between becomes uninterpretable. Anyone tracking capability over time is measuring against a ruler that changes length annually.

This is not primarily a measurement inconvenience. It is a governance problem. Regulatory frameworks, procurement standards, and safety commitments increasingly reference benchmark thresholds — and a threshold defined against a benchmark with a fourteen-month useful life is a threshold that expires before the rule implementing it takes effect. California's newly signed AI audit framework will have to confront this directly: independent verification organizations need stable evaluation instruments, and the field does not currently produce any.

What saturation does not mean

It does not mean AI does research-level mathematics autonomously. Tier 4 problems have known answers that can be checked. The generative act — deciding which question is worth asking, recognizing that a result matters, building the framework a field will use for twenty years — is not what Tier 4 tested.

The Navier-Stokes episode illustrated this precisely. Within hours of publication, the story became a priority dispute with NYU mathematician Tristan Buckmaster over the influence of unpublished work. That argument is about provenance and credit — human questions about who contributed the idea — on a result the machine formally verified. Verification was the easy half.

What to watch

FrontierMath Tier 4 v2 opening scores. If the successor benchmark opens at 30% rather than 5%, that tells you the capability is general. If it opens near 5% again, the saturation was substantially specific to this problem set.

Whether the shortcut problem gets measured. Epoch's observation about unintended shortcuts is currently an anecdote from problem authors. It should be a reported metric — the share of solutions that used the intended reasoning path. That number would be worth more than any leaderboard.

Hodge. OpenAI is rumoured to be working on another Millennium Prize problem. Fourteen months from 5% to 98% on research-level problem sets suggests the interesting constraint is no longer capability. It is which problems are formally checkable.

#frontiermath#epoch-ai#gpt-6-astra#benchmarks#math-reasoning

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.