The prize that built modern math—proving things—is now the line OpenAI can't afford to cross.
The Summary
- Mathematician Andreas Thom publicly accused OpenAI of using his pre-publication work to train models, raising questions about data provenance in mathematical AI
- This is the second researcher in days to challenge OpenAI's training data transparency around mathematical discoveries
- The core issue: if OpenAI scraped private ChatGPT interactions with mathematicians before announcing breakthroughs, it's not just unethical—it undermines the entire claim to discovery
The Signal
Mathematics is different from code or text. When you train on StackOverflow or Wikipedia, you're learning from published knowledge. When you train on unpublished mathematical work—proofs in progress, conjectures being tested, private conversations between researchers—you're not learning. You're stealing the answer key.
Andreas Thom's Mastodon posts suggest he and colleagues used ChatGPT to explore mathematical problems before OpenAI announced its mathematical reasoning capabilities. Now he's asking: did those interactions end up in the training data? If so, OpenAI didn't discover anything. It memorized homework it was never supposed to see.
This is the second mathematician in a week to raise this concern. The first dispute involved unpublished work that allegedly fed into OpenAI's models. The pattern emerging isn't about one angry researcher. It's about a systematic opacity problem at the exact moment OpenAI is claiming breakthroughs in formal reasoning.
"If your AI 'solves' a math problem it saw you solving last month, that's not reasoning. That's a parrot with a good memory."
The math community operates on proof. You show your work. You cite your sources. You don't claim credit for theorems you derived from someone else's unpublished notes. OpenAI is asking mathematicians to accept, on faith, that its models didn't benefit from private interactions on its own platform. That's not how math works.
Here's what makes this particularly thorny:
- ChatGPT's terms of service give OpenAI rights to user data for training
- Researchers often use AI tools to explore ideas before publication
- There's no audit trail showing which conversations fed which models
- Mathematical "discoveries" are worth billions in AI capabilities marketing
The implications go beyond hurt feelings. If leading AI labs can't prove their training data is clean, every mathematical result they publish is suspect. You can't build AGI on a foundation of "trust us." Math doesn't work that way. Science doesn't work that way.
OpenAI has been vague about its mathematical training corpus. That vagueness worked fine when the models were mediocre. Now that they're claiming frontier mathematical reasoning, the math community wants receipts. Not press releases. Receipts.
The Implication
This is a canary moment for AI transparency. If OpenAI can't or won't prove its mathematical models were trained on legitimate sources, every other claim about reasoning capabilities becomes suspect. The math community has the tools and motivation to expose contaminated training data. They will.
For anyone building AI systems: your training data provenance is now a liability, not just a technical detail. If you can't show clean lineage from source to model, especially in domains like math where proof matters, you're building on sand. The researchers are watching now.