OpenAI Dumped 722 Math Papers on GitHub. Three Are Already Gone.
An unreleased model attacked roughly 4,000 open problems at about three hours of ChatGPT Pro compute each, and OpenAI says nearly every result came from a single prompt to a single agent. A day later, a sign error took out three manuscripts.
A month ago, OpenAI's Navier-Stokes claim took a swarm of roughly 10,000 agents. On Tuesday evening, the company went the other direction: it posted 722 mathematics manuscripts to a public GitHub repository and told Scientific American that almost all of them came from a single prompt handed to a single AI agent.
The repository, openai/math, went up at about 6 p.m. EDT on October 6, according to Scientific American, under an Apache 2.0 license. The papers are grouped into 372 "families" of results, and OpenAI says they resolve or substantially advance open questions in mathematics and theoretical computer science. Scientific American reported that the batch includes a claimed solution to the four-dimensional Kakeya conjecture, improvements to important algorithms and progress toward the Riemann hypothesis. The model that produced them has not been released.
By Wednesday the count was 719.
How the batch was made
The repository's README is short on method but specific on scale. The "vast majority" of results came from the same fixed procedure using an unreleased internal OpenAI model. Over the course of the evaluation, the model was posed approximately 4,000 problems. On average, each result used three hours of ChatGPT Pro thinking compute. OpenAI then aggregated the output into families and manuscripts and applied what it calls "an appropriate level of significance" to decide what made the catalogue.
A family is not a single theorem. Per the README, it can bundle a principal result with companion arguments, consequences or alternative proofs, which is why 722 manuscripts map to 372 families. The 722 figure counts papers, not solved problems.
Two results sit outside the fixed procedure: work on a zero-free region for the Riemann zeta function, and a proof of the Hodge conjecture for CM abelian varieties. OpenAI says the write-up of the Re(s) > 11/12 zero-free region was also human-edited for readability. The company also published abridged reasoning summaries for ten families.
OpenAI framed the release as a byproduct. "As part of model development, we evaluate our models on open research problems," the README says, adding that the evaluations expanded "after performance on our existing mathematical evaluations saturated." Dan Roberts, an OpenAI research lead, told The New York Times that testing internal models helps build better tools, according to The Next Web.
Lean covers less than half
The Navier-Stokes release leaned on its Lean files as the argument. This one cannot, at least not yet. As of Wednesday's update, OpenAI's own history log puts formalization at 300 of 719 top-line results, about 42%. The remaining majority are natural-language manuscripts awaiting human review.
OpenAI is upfront about what that means. "Some of the unformalized results could have issues," the README says. "We will endeavor to fix any such issues quickly."
The first fixes arrived within a day. The repository's history file, dated October 7, records that a sign error in "Algebraicity of Weil classes on split abelian eightfolds" invalidated a cancellation argument and the construction used by two dependent papers. OpenAI withdrew all three, including a manuscript titled "The rational Hodge conjecture for products of K3 surfaces." It also revised 14 other manuscripts with proof repairs, corrected statements and clearer hypotheses, updated 13 more to cite the revised versions, and added six formalizations.
That is a quick, transparent correction cycle. It is also a preview of the review burden. Three withdrawals and 14 repairs in the first 24 hours, on a catalogue most of which no human has yet checked, tells mathematicians roughly what the next several months look like.
The disclosure fight
The release landed one week after a set of ground rules OpenAI helped prompt. On September 21, the company announced an independent advisory group of mathematicians to recommend how AI-generated results should be published, Scientific American reported. That group, the Advisory Group on Mathematics and Artificial Intelligence (AGMAI), hosted by the Institute for Advanced Study, issued guidelines on September 29.
Those guidelines asked labs to stop testing advanced mathematical problems on proprietary models, warned that internal models "risk creating a two-tier system where labs outrun the rest of the field," and asked that each released result come with the model, the prompts, a reasoning summary, the time spent and the estimated cost of computation.
OpenAI delivered part of that list. It disclosed average compute and some statistics, but no prompts, Scientific American reported, and the model remains internal. A spokesperson told the magazine that OpenAI is not bound by the group's recommendations but is trying to follow them, and that many of the new results are not yet understood by OpenAI's own mathematicians.
AGMAI put out its own statement the same night. Its advisory role, the group wrote, "should not be interpreted as a judgment of the impact of these results," and the release "is the beginning, not the completion, of the process of human understanding."
Receipts, and a field under load
The single-prompt claim is the part mathematicians are least willing to take on trust. MIT's Andrew Sutherland told Scientific American that claims about one-shotting problems with a single agent should be treated as unverified until the model is released and others can reproduce the results. "We should ask for receipts," he said. The OpenAI spokesperson also acknowledged that some results may have taken multiple attempts.
NYU's Tristan Buckmaster, who was at the center of last month's Navier-Stokes credit dispute, doubted OpenAI could have checked so many results at once. "I don't think they've done their sort of due diligence at all," he said, according to The Next Web.
Not everyone reads the dump as a threat. "To me, it's going to be a good thing for mathematics," the University of Toronto's Daniel Litt told Scientific American. The magazine also noted that Terence Tao has publicly criticized the frontier labs over the pace of their AI-generated results.
What actually changed
The Navier-Stokes result showed that a lab could throw enormous parallel compute at one famous problem. This release claims something different and, if it holds, more consequential: that a single agent with a three-hour budget now produces publishable-grade mathematics across hundreds of problems, and that the cost per result is closer to a long work session than a supercomputer run.
The bottleneck has moved. Generation is cheap. Verification is not, and 58% of this catalogue has no machine check behind it. Until the model ships and the prompts are public, the mathematics community is being asked to do the expensive half of the work on output it cannot reproduce. The sign error found on day one is the system working. It is also the reason nobody should count to 722 yet.
