Two separate disputes involving mathematicians and OpenAI have converged on the same uncomfortable question: what happens when researchers discuss unpublished work with an AI system that may later become capable of solving the same problems?
There is no public evidence that OpenAI copied a specific private ChatGPT or Codex conversation and turned it into a mathematical breakthrough. Neither Andreas Thom nor Tristan Buckmaster has proved that happened.
But their questions are more specific than the viral accusation that OpenAI simply "stole" their work.
They want to know whether nonpublic research shared through OpenAI products could have entered a training or model-improvement pipeline before later OpenAI systems produced closely related mathematical results.
OpenAI denies that specific user data was accessed while solving its Navier–Stokes result. The company has also acknowledged a narrower uncertainty: it says it cannot completely rule out the possibility that de-identified data derived from product usage helped improve its models.
That difference — direct access to a private conversation versus earlier information influencing a model — is the heart of the controversy.
It started with an AI result in group theory
On August 1, OpenAI published ten results in mathematics and theoretical computer science produced by an internal version of Astra.
Among them was a construction establishing the existence of a non-sofic group, addressing a question that had resisted mathematicians since sofic groups were introduced in the late 1990s.
The mathematical background is important because the AI result did not appear out of nowhere.
Earlier work by Gábor Kun, as well as joint work by Kun and Andreas Thom, provided key ingredients used in the eventual construction. The Alfréd Rényi Institute of Mathematics later highlighted the connection between their earlier results and OpenAI's proof.
Andreas Thom had been discussing related work with ChatGPT
Thom later published a three-part statement describing an exchange he had with OpenAI after the non-sofic-group announcement.
According to Thom, he and a colleague in Dresden had spent months discussing the expander matching problem and extensions of work involving Kun with ChatGPT.
After seeing OpenAI's result, he wanted to understand whether those conversations had any relationship to what the model had produced.
His question contained two distinct possibilities.
First, had the conversations entered training data or otherwise contributed to improving the model? Second, had the conversations been directly accessible to the system while it was solving the problem?
Thom says OpenAI researcher Mark Sellke responded, in full: Regarding your conversations with ChatGPT: that did not happen.
The answer initially appeared categorical. Thom later argued that it did not clearly establish which of his two questions had been answered.
Direct access and training are two different questions
This distinction matters.
An AI system could directly retrieve a user's private conversation while working on a problem. That would be one scenario.
A different scenario would involve information from previous product interactions contributing to an earlier model-training or improvement process. A later model could then use what it learned without retrieving the original conversation at all.
There is also a third possibility: independent discovery. A sufficiently capable model trained on published mathematical literature might reach an approach similar to one being explored privately by human researchers.
Thom argues that only OpenAI has the internal information needed to establish what happened in his case.
He also says he disabled model training on June 29. His point is that a setting changed on that date would not, by itself, explain how earlier conversations had been handled.
Thom has not published evidence showing that his conversations were actually included in Astra's training data.
Instead, his argument is about what has not yet been demonstrated publicly.
Why de-identification does not fully answer the research question
For ordinary personal data, de-identification is generally discussed as a way of separating information from the person who supplied it.
Research introduces another dimension.
An unpublished mathematical argument may remain intellectually valuable after the researcher's name, email address and account information are removed.
Thom's concern is therefore not simply whether OpenAI knew who supplied an idea. He is asking whether the intellectual content of nonpublic research could have contributed to model improvement.
Again, that is a question — not proof that it happened.
Thom's final post was far more critical
The third part of Thom's statement moved beyond technical questions about data handling.
He sharply criticized the response he says he received and also referenced the separate dispute involving Tristan Buckmaster and Levent Alpöge.
That qualification is important. Thom's criticism of individual OpenAI researchers is his own characterization of what happened.
It should not be treated as evidence that OpenAI stole unpublished work.
Then came the Tristan Buckmaster controversy
Thom's questions attracted more attention because another mathematician had just raised a strikingly similar issue.
Tristan Buckmaster and Levent Alpöge had been working on difficult finite-time blowup problems involving fluid equations. In his public statement, Buckmaster said the pair used several large language models throughout the project, including Anthropic's Claude and OpenAI's Codex.
He also made clear that the broader mathematical program predated their work, crediting earlier research by Diego Córdoba and Luis Martínez-Zoroa.
Buckmaster said their Codex sessions contained project drafts produced throughout the collaboration.
During later discussions with OpenAI, Buckmaster says he asked whether the internal model had been trained on, or had access to, those Codex sessions.
According to his account, he received a response saying the model did not look up user data. He then asked specifically about training and says that follow-up question was not answered at the time.
But one detail is often lost when the controversy is compressed into social-media posts: Buckmaster explicitly said he did not know whether his data had been used.
He did not present data theft as an established fact.
What OpenAI says happened in the Navier–Stokes project
OpenAI published a detailed account of its Navier–Stokes work on September 8.
The project involved a new internal model that OpenAI described as significantly more capable than GPT-6 Astra.
OpenAI says the effort began on September 1 after it heard rumors that progress had been made on two Millennium Prize problems.
Groups of AI agents were assigned different mathematical approaches. The group that ultimately produced the Navier–Stokes result involved roughly 10,000 concurrent agents.
The system reached its proposed solution on September 5, about 88 hours after the first agents were launched. OpenAI says formalization and verification in Lean took another 17 hours using GPT-6 Astra.
For the Navier–Stokes problem alone, the agents exchanged about 2.7 million messages and generated roughly 130 billion output tokens.
The proposed construction describes smooth fluid behavior that develops a finite-time singularity under smooth forcing, with trajectories spiraling inward as the flow stretches along its axis.
The sentence that keeps the controversy alive
OpenAI says its researchers and AI agents did not see Buckmaster and Alpöge's work before the researchers released it publicly.
The company also says no specific user data was accessed in order to solve the problem.
But OpenAI added a narrower caveat: it cannot completely rule out that de-identified data derived from the researchers' use of OpenAI products helped improve its models.
OpenAI described that possibility as unlikely.
That does not amount to an admission that Buckmaster's unpublished mathematics was used.
It does, however, explain why the debate did not end with a simple denial of direct data access.
There is also an important mathematical distinction. Buckmaster and Alpöge's public result concerned the forced Euler problem, while OpenAI's own Euler result concerned the unforced version. OpenAI says its proofs and the precise results are materially different.
The 100,000-researcher program changed the context
The controversy is arriving just as OpenAI is trying to put frontier AI tools in the hands of far more academics.
On July 29, OpenAI announced ChatGPT for Academic Researchers, initially selecting a 10,000-seat cohort and planning to expand access to 100,000 scientists, mathematicians and engineers through 2027.
The company says those dedicated research workspaces include business-grade privacy and security protections and that their data is not used to train OpenAI models by default.
That is an important distinction from a standard personal ChatGPT workspace.
The announcement also produced a provocative prediction on social media.
The post argued that researchers could eventually hand valuable academic material to AI companies, only for those companies to use much larger computing budgets to move faster than the original researchers.
After the Thom and Buckmaster controversies surfaced, that prediction began circulating again.
Its renewed popularity is understandable. But it remains speculation. OpenAI's published policy for the dedicated Academic Researchers workspace says its data is not used to train models by default.
What ChatGPT's data controls say now
For researchers using ordinary personal ChatGPT accounts, the details are different.
OpenAI's current Data Controls documentation says personal ChatGPT users can switch off Improve the model for everyone.
Once that setting is disabled, new conversations remain available in chat history but are not used to train ChatGPT.
OpenAI says Business, Enterprise, Edu and API customer inputs and outputs are not used to train its models by default.
The practical lesson is that researchers should not treat every OpenAI product as having identical data rules.
A personal ChatGPT workspace, an institutional environment, a dedicated academic workspace and a specialized tool can have different protections and controls.
What is actually established?
- Andreas Thom says he discussed related mathematics with ChatGPT before OpenAI announced its non-sofic-group result.
- He says he asked OpenAI about both direct access to those conversations and their possible role in model training.
- He believes the answer he received did not clearly resolve both questions.
- OpenAI's non-sofic result drew on earlier published mathematics, including work by Kun and joint work by Kun and Thom.
- Tristan Buckmaster says research drafts were present in Codex sessions during his collaboration with Levent Alpöge.
- Buckmaster says he asked OpenAI whether those sessions had been accessed or used for training.
- Buckmaster explicitly says he does not know whether his data was used.
- OpenAI says no specific user data was accessed to solve its Navier–Stokes problem.
- OpenAI says it cannot completely rule out an indirect contribution from de-identified product-usage data to model improvement.
- No public evidence currently proves that OpenAI copied a specific unpublished mathematical idea from a private ChatGPT or Codex conversation.
Why this matters even if no theft is ever proven
The larger story is not really about one proof.
AI systems are becoming research environments.
Scientists now use them to explore hypotheses, test unfinished arguments, inspect experimental results, write code and find weaknesses in ideas before publication.
A conversation with an AI assistant can therefore contain information that would once have lived only in a notebook, a private laboratory or an email exchange between collaborators.
That changes what "privacy" means.
If a researcher's name is removed from an unpublished idea, the idea itself may still be valuable.
And if researchers cannot tell how that intellectual content can move through an AI company's systems, some will hesitate before sharing their best unfinished work.
That would produce an awkward paradox: the more useful AI becomes for frontier research, the more careful researchers may need to become about what they give it.
Should researchers stop using ChatGPT for unpublished work?
Not necessarily.
The more useful response is to understand the environment being used.
Before putting unpublished research, confidential datasets or commercially sensitive ideas into an AI tool, researchers should check the current data policy for that specific product and workspace.
For personal ChatGPT accounts, that includes reviewing the Improve the model for everyone setting.
For institutional or collaborative research, a managed workspace with explicit privacy protections may make more sense than a standard personal account.
And for work whose confidentiality is critical, the data policy should be treated with the same seriousness as access to an unpublished manuscript or private research repository.
So, did OpenAI use private ChatGPT data?
The evidence available publicly does not provide a definitive yes.
OpenAI denies directly accessing specific user data in its Navier–Stokes research effort.
Thom and Buckmaster have raised a different question: whether earlier interactions with OpenAI products could have contributed to model training or improvement before those models produced related mathematical work.
OpenAI's own statement leaves a narrow uncertainty around de-identified usage data, while saying that possibility is unlikely.
That is not proof that private research was taken.
It is also why the controversy cannot be reduced to a simple claim that nothing remains to be explained.
For now, the most defensible conclusion is straightforward: the allegation of research theft remains unproven, while the debate over training-data transparency remains unresolved.
Sources and further reading
- OpenAI — Ten advances in mathematics and theoretical computer science, August 1, 2026.
- OpenAI — On the Navier–Stokes Millennium Prize Problem, September 8, 2026.
- Andreas Thom — Three-part public Mathstodon statement on OpenAI and unpublished mathematics.
- Tristan Buckmaster — Public statement concerning his collaboration with Levent Alpöge and their use of AI tools.
- Alfréd Rényi Institute of Mathematics — Background on the non-sofic-group breakthrough and the role of previous Kun and Kun–Thom research.
- OpenAI — ChatGPT for Academic Researchers.
- OpenAI Help Center — Data Controls FAQ.
Cybero Plus will update this article if OpenAI or the researchers involved publish additional evidence, corrections or substantive responses.