OpenAI Math Training Data: Mathematicians Demand Proof

Two mathematicians in one week now demand proof OpenAI didn't use their unpublished work — and OpenAI's own blog post concedes it can't provide any. Andreas Thom (non-sofic groups) told The Verge today that after asking researchers Sébastien Bubeck and Mark Sellke directly whether his ChatGPT conversations fed the model's celebrated result, "no such qualification, explanation, or evidence was given." His verdict: "dishonesty to say the least."

The receipt

The smoking gun is OpenAI's own wording in its Navier–Stokes announcement: "no specific user data was accessed" — followed immediately by "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models." Thom's rejoinder is the one that matters: "De-identification may remove a name; it does not remove the intellectual content of a mathematical idea." Only OpenAI holds the training-pipeline data needed to verify the denial. The burden of proof is on the party that can't prove it.

Why it matters

Read this next to Terence Tao's post from last night warning that premature AI exposure to open problems flattens research incentives, and the pattern is complete: mathematicians are now assuming their casual model interactions are inputs to a competitor. Researchers told The Verge this will push the field secretive — if a rumor of progress can ignite a race with a well-resourced lab, you stay quiet. That's a collective-action problem, and it's already happening.

The fix everyone's circling is provenance: independent attestation of what entered a training run. We wrote about trajectory watermarking yesterday as an academic idea. It just became a governance demand. The labs that can produce a receipts chain for their training data will be the ones mathematicians — and every other expert community with pre-publication IP — actually talk to. "Only OpenAI has the relevant data for that" is not a defense. It's the indictment.

Related