AI

Fields Medalists vs OpenAI's 722 AI Math Manuscripts

OpenAI posed ~4,000 problems to get 372 result families and published 722 manuscripts, eight days after an IAS advisory group asked labs to stop testing advanced mathematics on proprietary models.

Filed by
Published
Read time7 minutes
Fields Medalists vs OpenAI's 722 AI Math Manuscripts

Three weeks after twenty-five Fields Medalists signed a declaration asking AI labs to change how they do mathematics, OpenAI published 722 mathematical manuscripts produced by a model nobody outside the company can run.

The release went up on October 6, 2026, as a GitHub repository under an Apache-2.0 licence, with Lean proof files, ten abridged reasoning summaries, and a README describing the whole thing as a by-product of model evaluation. Sam Altman's framing on X the following morning was "We are entering a new era of discovery now." The mathematicians who signed in September had used a different word for it: misalignment.

Both sides are describing the same artefact. What separates them is a question of what a mathematical result is for, and the repository happens to contain enough data to put numbers on the disagreement.

The funnel, in the lab's own figures

The README states the production terms openly. The model was posed approximately 4,000 problems. The output was consolidated into 722 manuscripts covering 372 result families, where a family groups a principal result with its companion arguments, consequences and alternative proofs. The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking.

Read as a pipeline, those numbers say something the announcement does not.

Horizontal bar chart showing four stages: approximately 4,000 problems posed, 722 manuscripts published, 372 result families, and 162 main results formalized in Lean.

Sources: openai/math README for attempts, manuscripts and families; lean/formalization.yaml for the formalized count, read October 7, 2026.

Stage Count Share of the 4,000 attempts
Problems posed to the model ~4,000 100%
Manuscripts published 722 18.1%
Result families 372 9.3%
Main results with a Lean formalization 162 4.1%

Roughly 4,000 attempts produced 372 result families — a 9.3% yield — and 162 of those results, 4.1% of the original attempts, carry a machine-checked proof. The 162 figure is not in the announcement; it comes from counting the entries in the repository's own formalization catalogue, which lists 162 unique manuscripts against the 722 linked from the manuscript map.

That is a respectable hit rate for open research problems, and it is also an industrial one. Nine percent of attempts converting is a throughput statistic, the kind of number a process engineer quotes. It is not a number any individual mathematician has ever been in a position to report, because no individual has ever been able to attempt four thousand open problems.

What it cost, as far as anyone can tell

OpenAI priced the work in an unusual unit: ChatGPT Pro thinking time. At roughly three hours per result, the 722 published manuscripts embody on the order of 2,166 hours of Pro-grade inference — about ninety days of continuous thinking if it had run on one stream instead of in parallel. The roughly 3,300 attempts that did not make the cut are not priced in the disclosure at all, so the true compute behind the release is higher than that figure by an unknown multiple.

The dates add the other half of the picture. Of the 722 manuscript folders, 721 carry a date in the folder name, and they cluster into fifteen days: September 22 through October 6. That is an average of 48 research manuscripts per day, with a peak of 193 dated September 24 and a second wave of 112 dated October 5, the day before publication.

For comparison, the arXiv mathematics section receives on the order of a hundred submissions a day from the entire world. One unreleased model, in fifteen days, produced a body of claimed results of comparable daily volume.

What a "family" is, and why the count is softer than it looks

The gap between 722 manuscripts and 372 families is not padding, but it does mean the two figures answer different questions and should not be used interchangeably.

A family, as the README defines it, groups related papers: a principal result plus companion arguments, consequences, or alternative proofs of the same thing. So 722 is the number of PDFs, 372 is roughly the number of distinct mathematical claims, and the ratio — 1.94 manuscripts per family — says the collection averages a principal paper and one supporting paper per result. Anyone quoting "722 new theorems" is double-counting by about half. Anyone quoting 372 is closer, with the caveat that the significance threshold for inclusion was set internally by the lab rather than by a referee.

The repository is also more carefully assembled than a dump. Every manuscript directory carries a BibTeX block so results can be cited rather than screenshot, the licence is Apache-2.0, and OpenAI has committed to preserving the public release history so corrections appear as new versions with the earlier ones still reachable. Two results are flagged as departures from the standard procedure: the zero-free region for the Riemann zeta function, whose writeup was human-edited for readability, and the proof of the Hodge Conjecture for CM abelian varieties. Ten abridged reasoning summaries cover a further selection, including the irrationality exponent of π, the symmetric and general Mahler conjectures, and Kaplansky's direct-finiteness conjecture in characteristic two.

Those are the affordances of a group that expects to be audited, which makes the one withheld affordance — the model — stand out more, not less.

The declaration this landed on top of

The timing is what makes this a governance story rather than a benchmark story.

On September 8, OpenAI announced that an internal system had produced a proposed solution to the Navier–Stokes problem. Three days later, twenty-five Fields Medalists — including Terence Tao, Peter Scholze, Maryna Viazovska, Pierre Deligne, Manjul Bhargava and Cédric Villani, spanning medals from 1978 to 2026 — published "A Severe Misalignment of AI in Mathematics", which has since collected thousands of additional signatures.

Their objection was precise, and it was not that the proofs were wrong. Problem-solving, the declaration argues, is a proxy for the real goal of shared conceptual understanding. Optimise the proxy hard enough and you invert the relationship: the field accumulates answers faster than it accumulates comprehension, and the attribution and priority conventions that mathematics runs on stop working. That last concern was not hypothetical. It later emerged that the Navier–Stokes result had landed on top of work already under way by another team of mathematicians who were also using AI.

The Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study then ran a consultation, received more than 600 replies, and published its recommendations on September 29. One of them reads: "At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community… we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models."

Eight days later OpenAI published 722 manuscripts from a proprietary model that remains inaccessible to the broader scientific community, and cited the advisory group's advice in the announcement.

The group asked labs to stop testing advanced mathematics on models nobody else can run; OpenAI thanked them for the advice and shipped 722 results from exactly such a model. Asked about the gap, OpenAI told Gizmodo that the group's "advice and public recommendations have informed how we're sharing the results," and that it would keep incorporating community feedback.

That is, read carefully, an answer about sharing, not about testing. OpenAI adopted the publication protocol — citation blocks, revision history, Lean files, reasoning summaries — and declined the part that would have required releasing the model. The company does say it is "working to responsibly release" it.

Where this leaves a working mathematician

The practical position is awkward. A researcher in, say, spin glasses now has to decide whether to read the Mézard–Parisi family in the repository, and there is no referee report to lean on, no author to email, and — for roughly four out of five manuscripts — no machine-checked proof either. Checking one by hand costs weeks. Ignoring it risks publishing something the corpus already contains.

OpenAI's stated remedy is to fund workshops, conferences and special programmes to build understanding around the major results, which is an admission that the bottleneck has moved. Producing the mathematics turned out to be the cheap part; metabolising it into something a seminar room can follow is now the expensive one.

That gap is where a surprising amount of ordinary research tooling has migrated. Reading groups, explainer threads, summary decks and talk slides are all doing distribution work that used to be handled by the slow pipeline of journals and conferences, and AI tools have followed the work there rather than into the proofs themselves. ChatSlide, an AI slide generator that turns a paper or a set of findings into a presentation, sits in that second category: it makes no claim on correctness, and it addresses the step the Fields Medalists were actually worried about, which is whether anyone ends up understanding the thing.

None of which settles the argument. OpenAI has demonstrated that a frontier model can run mathematics as a throughput process at roughly nine percent yield. The mathematical community has said, with unusual unanimity for a field that agrees on very little, that throughput is not what it is optimising. The repository is the first artefact where both claims are true at the same time, and it is public, dated, and counted.

Counts in this article were derived from the public openai/math repository on October 7, 2026: 722 manuscript paths in CONTENTS.md, 162 unique manuscript ids in lean/formalization.yaml, and folder-name dates for the 721 manuscripts that carry one.

Disclosure: ChatSlide is operated by a company that also publishes Tech Forum.

About the author

Steph Moreau

Steph Moreau is a senior correspondent at Tech Forum, specializing in fintech, enterprise software, and venture capital. Her sharp analysis of funding rounds and market trends helps readers navigate Canada's evolving tech economy.

Keep reading

More from Tech Forum