SDSignal Desk

Mathematicians want proof OpenAI didn’t use their work

Sep 10, 2026, 4:00 AM · The Verge

Image: The Verge

After OpenAI’s math breakthroughs, Andreas Thom joins Tristan Buckmaster in demanding proof that chatbot interactions didn’t quietly train the models that later raced the field.

Why it matters

A second mathematician is publicly challenging OpenAI over training-data provenance after the company’s recent math announcements. Andreas Thom says interactions he and colleagues had with ChatGPT before OpenAI’s results may have contributed — and that OpenAI’s answers haven’t closed the question.

One of ten results OpenAI highlighted involved non-sofic groups, Thom’s area; OpenAI later acknowledged heavy dependence on prior work by Thom and Gábor Kun after criticism for thin credit. Thom wrote researchers Sébastien Bubeck and Mark Sellke asking whether his ChatGPT conversations were in training data or available to the reasoning process.

He says the reply addressed direct access to conversations, not whether material entered training corpora used to improve models — a distinction he calls dishonest. OpenAI did not immediately comment to The Verge.

From the desk

This is a high-stakes trust story, not a footnotes spat. If frontier labs use product traffic — even de-identified — to sharpen models that then compete with the same experts for priority, the research commons breaks. Thom’s line lands: de-identification can strip a name without stripping an idea.

We’re not asserting OpenAI trained on Thom’s chats; we don’t have that evidence, and neither does he. That’s the point. Only OpenAI can prove the negative, and mathematicians aren’t positioned to reverse-engineer the pipeline. After the Millennium Prize / Navier-Stokes episode — where OpenAI denied accessing Buckmaster and Levent Alpöge’s specific work, yet would not fully rule out de-identified product-derived improvement — the same narrow wording is now a pattern people notice.

Useful AI in science needs provenance discipline equal to the brag. Credit amendments after backlash are better than silence, still worse than getting the literature and the data story right on day one. The harm if this scales is cultural: mathematicians told The Verge they fear rumors of progress will trigger well-funded lab races, pushing the field toward secrecy instead of open discussion.

I’m watching for a concrete disclosure standard — what user-derived data can and cannot influence research models — with auditable boundaries, not another carefully lawyered “while unlikely.” Without that, every future OpenAI math win will arrive under a cloud of process doubt even when the math is right.

Context

The episode follows OpenAI’s pursuit of a Millennium Prize problem after hearing online that others were close. Verification of the claimed solution remains a separate scientific process; the dispute here is about data ethics, credit, and competitive norms around unpublished ideas.

Who feels it

Research mathematicians
Incentive to withhold drafts and chatbot explorations if product use might feed rival model improvement.
OpenAI and frontier labs
Pressure to publish harder guarantees on training exclusions for research-sensitive product data.
Scientific publishers and societies
May need clearer norms for AI-assisted priority, citation, and disclosure when labs and academics overlap.

What to watch

  1. Whether OpenAI issues a detailed public accounting of math-related training and user-data boundaries.
  2. Follow-up from Thom, Buckmaster, and others if more correspondence surfaces.
  3. Signs of increased secrecy or pre-print withholding in core math communities.

Read the original

Continue at the source.

The Verge

Companies: OpenAI

Also covering this