AERIOXFLUX
Frontier Labs
Frontier Labs · openai

Astra Closed Ten Open Problems and Left Receipts in Lean

OpenAI's unreleased next model produced machine-checkable proofs for ten open questions in math and theoretical CS — and the $2,000 token bill is the part the field will argue about.

Flux Desk·2026-08-07·5 min read

On August 1, OpenAI announced that an internal version of Astra — its next major model, still unreleased — had produced solutions to ten open problems across mathematics and theoretical computer science. It published a 249-page manuscript, model reasoning walkthroughs, and the thing that actually matters: Lean 4.32.0 certificate files, built against mathlib, in a public repository under Apache-2.0.

The repo went up at 06:10 UTC the same morning.

Every previous "AI does mathematics" announcement has been an argument about evaluation. This one isn't, and that is the whole story.

What's in the box

The ten results are not a themed set. They span group theory, operator algebras, high-dimensional geometry, complexity, lattice cryptography, and extremal combinatorics:

  • A construction proving the existence of non-sofic groups — a central open question in group theory.
  • A disproof of Connes's rigidity conjecture.
  • An improved sphere-packing bound in high dimensions — the first movement on that bound since 1978.
  • An n⁴/log n lower bound for computing the permanent.
  • An n^(1/400)-factor hardness result for the Euclidean closest vector problem.
  • An exponential parallel repetition theorem for two-player quantum games.
  • Resolutions of Erdős problems 146, 180, and 183.
  • Improved bounds for binary and spherical codes.
  • A solution to Ehrhart's volume conjecture.

Thomas Bloom, the University of Manchester mathematician who maintains erdosproblems.com — the registry that tracks exactly this kind of claim — called the results big news. Noam Brown, on OpenAI's side, was careful to note there are no Millennium Prize Problems in the set. Yet, he added.

No outside mathematician has publicly rejected any of the ten.

Why Lean changes the argument

The reason prior AI-math claims collapsed into methodology fights is that natural-language proofs require a human referee, and referees are slow, scarce, and inclined to give a machine less benefit of the doubt than a colleague. A model that produces a plausible eight-page argument has produced a homework problem for an expert.

A Lean certificate is a different object. It compiles or it doesn't. The theorem prover checks every inference step against mathlib's formalized foundations, and the check is mechanical, reproducible, and cheap. Anyone with Lake installed can clone the repo and run it tonight.

That collapses the trust question down to one narrow, honest residue: does the formal statement in the file say what the English claim says it says? Formalization is a translation, and translations can be wrong or weaker than advertised. That is a real objection, it is the objection specialists will actually raise, and it is checkable by reading a few lines rather than by refereeing a manuscript for six months.

OpenAI also disclosed that humans prepared the manuscripts before formalization. Read that carefully — it means the write-up and the framing are collaborative, and the load-bearing claim is narrower than "the model wrote the paper." It's "the model found arguments that survive mechanical verification." That's still enormous. It's just a different sentence.

The $2,000

OpenAI put the token cost of finding all ten solutions at roughly $2,000 at Sol API rates.

That number is doing more work than any of the ten theorems.

Frontier mathematics has been, economically, an infinite-cost activity: the input is a scarce human who has spent fifteen years acquiring the taste to know which questions are tractable, and the output arrives on a schedule nobody controls. There has never been a price tag on a sphere-packing improvement because there has never been a supply curve.

Two thousand dollars is a supply curve. It implies that the binding constraint on a certain class of result is no longer insight but problem selection — knowing what to point the thing at. If that holds even partially, the scarce resource in mathematics shifts from proving to curating, and the people who benefit most are the ones maintaining registries of well-posed open questions. Bloom's site is suddenly infrastructure.

Two caveats keep this from being a clean number. It is an inference-cost estimate on an internal model at published API rates, which is a hypothetical price, not a receipt. And it excludes every failed attempt — search costs are only cheap if you count the searches that worked.

What it doesn't answer

The model is private. There is no public interface, no release date, no external testing. The methods, training data, and search procedure are undisclosed. What the world can verify is the output, and only the output.

That asymmetry is exactly what the Leiden Declaration was drafted for — the push for stricter disclosure standards around AI-assisted results. OpenAI's release is unusually generous by the standards of proprietary labs and still falls short of what reproducibility would require: you can confirm the theorems are true, and you cannot confirm anything about how they were found, or whether a second run finds them again.

Peer review also remains a separate process. Lean says the argument is valid. Journals decide whether it is interesting, correctly attributed, and situated in the literature — and those judgments are not automatable.

What to watch

Whether specialists ratify the formalizations. The first substantive critique will be about a statement in a .lean file being weaker than the English abstract. If none materializes in six weeks, the results are effectively banked.

Whether other labs answer in Lean. Google DeepMind has shipped competitive math systems for years without publishing certificate files at this granularity. Formal verification is now the format for making a claim in this space; the lab that keeps announcing in prose is choosing to be doubted.

Whether Astra's release keeps this capability. Internal versions are not products. The gap between "an internal checkpoint did this" and "you can do this" has historically been months and a lot of safety-driven sanding.

The read

The ten theorems are the headline and the receipts are the story. By shipping machine-checkable certificates instead of a press release with a plot, OpenAI moved the burden of proof to where it can actually be discharged — and made every future claim that arrives without them look like an argument someone is trying to win rather than a result someone is trying to establish.

The mathematics is real. The verification standard is the news.

#openai#astra#lean#formal-proofs#mathematics

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.