
Can Mathematicians Still Trust OpenAI With Unpublished Proofs?
A Mathstodon post revives the FrontierMath funding controversy, showing mathematicians keep learning OpenAI's benchmark terms only after handing over unpublished work.
Can a mathematician still hand OpenAI an unpublished proof without regretting it? Researchers are asking that question again, and the timeline shows why the answer keeps getting worse instead of better.
November 2024: Epoch AI launches FrontierMath, a set of unpublished, expert-level math problems built specifically so no model could have memorized the answers from its training data. Research mathematicians are recruited to contribute problems on the understanding that the set stays genuinely unseen by any lab.
December 2024: mathematicians learn that OpenAI funded FrontierMath and held early access to the problem set and solutions, an arrangement Epoch AI had not disclosed to the contributors who had just handed over unpublished work. The disclosure surfaces only after outside pressure, not as a voluntary heads-up.
Days later: Andreas Thom posts on Mathstodon, asking the exact question those contributors were implicitly answering back in November, except now it’s out loud and in public. The December disclosure did not settle whether the arrangement was a one-off mistake or a pattern, and Thom’s post is the field saying so.
The through-line isn’t that OpenAI lied about anything specific. It’s that mathematicians keep discovering the terms of these arrangements after they’ve already contributed the work, never before. Every round of this story has the same shape: a lab gets access to unpublished material through a benchmark or research partnership, the funding relationship surfaces later than the contribution did, and the field has to decide, retroactively, whether that was acceptable. That’s a structural problem, not a lapse, and it will keep recurring as long as benchmark funding and model training sit inside the same company deciding what counts as disclosure.
If you’re a researcher weighing whether to share an unsolved problem with an AI lab this year, the practical answer is: ask who funds the benchmark before you ask how good the model is. The same skepticism belongs anywhere a lab grades its own homework — we’ve covered why AI models still stumble on math tests that look solved and how much a builder should actually trust an AI’s own quality reports.
Expect at least one more math-benchmark funding arrangement to become public under pressure, rather than by voluntary disclosure, before the middle of 2025 — the incentive that produced FrontierMath’s quiet arrangement hasn’t gone anywhere.
One AI signal a day. 90 seconds. No fluff. Get the next one in your inbox before it breaks — subscribe at /subscribe/.