GovernmentAI-TechBusinessScienceSportsEntertainmentGeneral
AI-Tech

OpenAI faces fresh research dispute over mathematics data

Mathematician Andreas Thom says the company failed to prove private chatbot conversations did not feed an automated proof of non-sofic groups.

​openai
Source: OpenAi LLC (Public domain)
Published10 Sep 2026, 12:52 Last updated10 Sep 2026, 12:52 Sources
Show reference links Marks each sentence drawn from a source or a contributor

A second prominent mathematician has publicly challenged OpenAI over the origins of data behind its automated theorem discoveries, accusing the company of evasive behavior and demanding proof that private user prompts were not absorbed into its training pipeline.1

Andreas Thom, a mathematics professor at Technische Universität Dresden, voiced the challenge following OpenAI's announcement of 10 automated results produced by an internal version of its forthcoming Astra model.12 One of those results solved a question open for more than 25 years by constructing a non-sofic group.21 OpenAI acknowledged that the mathematical construction relied directly on previous theorems proved by Thom and Gábor Kun, a senior research fellow at the HUN-REN Alfréd Rényi Institute of Mathematics in Budapest.12 Thom stated that his own interactions with the ChatGPT chatbot before the announcement touched on those precise methods, prompting questions about whether private technical prompts helped train the systems that solved the problem.1

The mechanics of the non-sofic proof

The mathematical dispute centers on infinite algebraic objects known as sofic groups, a concept introduced by Mikhail Gromov in 1999.21 Benjamin Weiss subsequently asked whether every countable group is sofic, framing a problem that remained unsolved for more than two decades.2 In technical terms, a group is called sofic if every finite fragment of its multiplication table can be approximated arbitrarily accurately by permutations of a finite set.2 The class contains all amenable and residually finite groups.2 Earlier research by Gábor Elek and Endre Szabó at the Rényi Institute established that sofic groups satisfy Kaplansky's Direct Finiteness Conjecture and Lück's Determinant Conjecture, and proved that every sofic group is hyperlinear.2

OpenAI's Astra model generated an argument establishing the existence of a non-sofic group, verifying each step inside the Lean proof assistant. According to an analysis published by the HUN-REN Hungarian Research Network, the machine proof rested on two central theoretical pillars developed by Kun and Thom.2 Kun had previously proved that finite graphs approximating a group with Kazhdan's property (T) can be decomposed into uniformly expanding components after negligible modification.2 In separate joint work, Kun and Thom demonstrated that sufficiently good approximations supported on a single expander impose strict finite-approximability constraints on groups commuting with the property-(T) action.2

The Astra system developed a procedure to match Kun's separate expander components, isolating an expanding approximation compatible with the Kun-Thom theorem.2 The system then showed that an embedded copy of Thompson's group V violated the required finite-approximability condition, yielding a contradiction that confirmed the group was non-sofic.2 OpenAI's manuscript explicitly cited Kun's decomposition theorem and the Kun-Thom result as the primary mathematical inputs to that section of the proof.2

Disputed communications and training pipelines

OpenAI initially published its writeup without fully crediting recent contributions from Thom and Kun, drawing criticism across mathematical circles before the company amended the text.1 Thom reported that he became suspicious after reviewing the text because Astra exhibited detailed mastery of analytical techniques that were neither obvious nor widely viewed as the most promising path forward.1

Thom contacted OpenAI researchers Sébastien Bubeck and Mark Sellke, who also works as a statistician at Harvard University, inquiring whether his private ChatGPT interactions had been incorporated into the model's training data or exposed to its reasoning processes.1 Sellke replied to Thom, but Thom concluded the response addressed only whether engineers had directly inspected his chat logs during the project, rather than clarifying whether user logs were included in broader pre-training or fine-tuning datasets.1

Thom criticized that distinction on Mastodon, writing that no explanation or evidence was provided regarding training ingestion and calling the explanation dishonest to say the least.1 Thom stated that Sellke's categorical answer was unjustifiably broad and materially misleading.1 Thom noted that outside researchers lack the computational access and technical logs needed to audit OpenAI's internal training pipelines independently, arguing that the burden of documentation rests entirely on the company.1

Parallel concerns over competition and credit

The dispute echoes an earlier challenge raised by Tristan Buckmaster, a mathematics professor at New York University.1 Buckmaster had questioned whether OpenAI's systems gained an advantage from his work inside OpenAI's Codex environment while he collaborated with Anthropic researcher Levent Alpöge on the Navier-Stokes equations, one of the Clay Mathematics Institute's Millennium Prize problems.1

OpenAI had pursued the Navier-Stokes equations after monitoring online rumors that external academic researchers were approaching a breakthrough.1 In its public statement on that work, OpenAI denied that its researchers or autonomous agents viewed user work prior to public dissemination, maintaining that specific user data was not accessed to solve the equation.1 The company added a qualification acknowledging it could not completely rule out that de-identified data derived from product usage helped improve underlying models.1 Thom rejected that distinction, stating that de-identification removes account names while leaving the mathematical concepts intact.1 Thom added that utilizing nonpublic research supplied by users to compete against those same users in academic publication without consent or disclosure represents an indefensible practice.1

Beyond group theory, OpenAI reported that Astra completed proofs spanning sphere packing, coding theory, operator algebras, circuit complexity, quantum complexity, lattice problems, discrete geometry, and extremal combinatorics.2 For academic researchers who regularly test commercial language models on unsolved technical problems, the absence of clear data governance boundaries has created hesitation, with researchers warning that competitive data practices will encourage greater secrecy throughout the scientific community.1

OpenAI did not immediately provide comment on the matter.1

Reporting note: this piece draws on reporting from The Verge published September 10, 2026, and technical documentation from the HUN-REN Alfréd Rényi Institute of Mathematics published August 5, 2026. Primary verified the mathematical history and citations directly against the institutional records.

Source: The Verge, September 10, 2026.

References

This article is based on 2 sources, listed in the order they are cited.

  1. 1 TV The Verge announcement · 10 Sep 2026 Mathematicians want proof OpenAI didn’t use their work See the source
  2. 2 HM HUN-REN Magyar Kutatási Hálózat third party · 5 Aug 2026 OpenAI See the source