Mathematicians Raise Transparency Concerns Over OpenAI's Math Training Data
A second mathematician has joined growing concerns about OpenAI's data practices in the development of mathematical AI models. Andreas Thom, a mathematician, posted on Mastodon questioning whether interactions his colleagues had with ChatGPT before OpenAI's public announcement may have inappropriately contributed to the model's mathematical capabilities.
Thom joins an earlier critic who raised questions about whether OpenAI benefited from unpublished mathematical research. The core issue centers on transparency: researchers are demanding clarity about what data was used to train models that have demonstrated increasingly sophisticated mathematical reasoning.
The accusations label OpenAI's practices as "dishonest" and ethically problematic. Researchers argue that the mathematical community deserves to know whether their prior work—whether published or not—was incorporated into training datasets without acknowledgment or consent.
This controversy highlights an ongoing tension in the AI field regarding training data transparency. As language models demonstrate capabilities in specialized domains like advanced mathematics, the question of how those capabilities developed becomes increasingly important to the researchers whose work may have influenced them.