SDSignal Desk

OpenAI’s math solutions aren’t meeting the field’s standards yet

Oct 8, 2026, 11:10 AM · TechCrunch

Image: TechCrunch

OpenAI says its models cracked hundreds of hard math problems, but a proof nobody understands is not finished mathematics, and the field's own guardrails were only partly followed.

Why it matters

This week OpenAI released hundreds of claimed solutions to difficult open math problems and said it had consulted an advisory group of leading mathematicians to avoid a repeat of past controversy. TechCrunch reports the release fell short of that group's standards, especially on the point mathematicians care about most: whether humans actually understand the result.

The advisory group, hosted by the Institute for Advanced Study, published guidelines at the end of September. Its first request was to stop testing advanced problems on proprietary models, which OpenAI's release says it is doing. Only 10 of the 719 manuscripts included the model's chain of thought. And a new paper from mathematicians at the University of Cambridge and King's College London documents at least two discrepancies between the plain-language proof and the formal Lean code behind OpenAI's solution to a problem derived from the Navier-Stokes equations.

From the desk

We want to be fair to the ambition here. If AI systems can produce real progress on open problems, that is one of the most hopeful things this technology could do. Mathematics is also one of the few places where claims can, in principle, be checked mechanically. That is why this matters so much: if the checking breaks down in math, it breaks down everywhere.

And the checking is where the cracks are. The pitch for formal proof languages like Lean is that once a proof compiles, it is correct. But that only holds if the formal statement says what the human-language claim says. The Cambridge and King's paper is about exactly that gap, the translation step. The discrepancies it found do not disprove OpenAI's solution, as TechCrunch notes. They do mean a compiled proof can quietly be a proof of something slightly different from the advertised result. The advisory group asked for machine-readable metadata linking the plain-language and formal versions so others could audit that link. OpenAI did not include it.

The deeper objection is cultural, and I think it is right. When a person proves something, they own it: they write it up, give talks, answer questions and teach the method to others. Terence Tao's complaint is that AI prompters move on once a target is 'solved' and cannot engage with the result. Harvard's Melanie Wood put the timing bluntly to TechCrunch: at release there is no human understanding yet, and the work begins then. The advisory group's suggestion that OpenAI help fund the mathematicians who will have to digest these results is not a courtesy. It is the cost of doing this responsibly.

Where this leads if it scales worries us. Dumping hundreds of manuscripts at a time turns a small, careful community into an unpaid review queue for a well-funded lab, with the headline credit landing before anyone has checked the work. That pattern will not stay in math. It is the same shape as AI-written code, research and legal filings outrunning the people meant to verify them.

Our read: the results may well contain real advances, and we hope they do. But a claim is not a result until a community understands and accepts it. OpenAI asked for expert guidance; the credible move now is to follow the parts it skipped.

Context

The Advisory Group on Mathematics and Artificial Intelligence, or AGMAI, is made up of nine prominent researchers at institutions around the world and is hosted by the Institute for Advanced Study, which TechCrunch notes is not affiliated with Princeton University. In a statement on the release, the group said it is up to the mathematical community to judge how well its recommendations were followed, and it did not respond to TechCrunch's request for a fuller evaluation.

Who feels it

Mathematicians
Hundreds of manuscripts land on the community at once, with limited reasoning traces and no formal-to-informal mapping to speed review.
AI labs
The advisory group's guidelines are becoming the yardstick for credible math claims, and partial compliance will draw public scrutiny.
Formal verification researchers
The translation gap between natural-language proofs and Lean code is now a front-line research and tooling problem.
Funders and institutions
Pressure is building for labs to pay for the human review their releases require.

What to watch

  1. Whether OpenAI publishes metadata linking its plain-language proofs to the Lean code
  2. Expert review of the Navier-Stokes-derived solution in light of the new paper
  3. A fuller public assessment from the advisory group
  4. Whether OpenAI commits funding for mathematicians to evaluate the results
  5. How other labs handle math releases against the same guidelines

Read the original

Continue at the source.

TechCrunch

Companies: OpenAI