OpenAI Claims Solved Navier-Stokes Math Problem, Drawing Scrutiny Over Training Data
External researchers question whether their unpublished draft work in Codex contributed to the company's Millennium Prize breakthrough.
Key highlights 路 2 min read
- OpenAI claims to have solved the Navier-Stokes existence and smoothness problem, a legendary 90-year-old challenge in fluid dynamics that stands as one of mathematics' seven Millennium Prize Problems.
- The announcement, first covered in detail by The Verge, immediately sparked a dispute over data provenance.
- Buckmaster raised alarm over the similarity in mathematical approaches, questioning whether OpenAI tapped private Codex sessions to fuel its internal system.
The Scale ReportOpenAI claims to have solved the Navier-Stokes existence and smoothness problem, a legendary 90-year-old challenge in fluid dynamics that stands as one of mathematics' seven Millennium Prize Problems. The company announced the finding on Tuesday, stating it used an unreleased internal model running alongside 10,000 concurrent agents to produce the proof.
The announcement, first covered in detail by The Verge, immediately sparked a dispute over data provenance. Just one day before OpenAI went public, Tristan Buckmaster, a mathematics professor at New York University, published related work completed alongside Levent Alp枚ge, a researcher at Anthropic. Buckmaster noted that his team had drafted their mathematical steps inside OpenAI's Codex coding environment and Anthropic's Claude before OpenAI revealed its own solution.
Data Access Questions
Buckmaster raised alarm over the similarity in mathematical approaches, questioning whether OpenAI tapped private Codex sessions to fuel its internal system. "I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project," Buckmaster stated. "I was told the model did not look up user data. I asked again, about training, and I did not get an answer."
OpenAI attempted to defuse the controversy by maintaining that "no specific user data was accessed in order to solve this problem." However, the lab conceded a caveat, adding that "while unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."
Sebastien Bubeck, a member of technical staff at OpenAI, defended the company's autonomy on social media. "We did not see any of their work until they released it publicly last night," Bubeck said. "One can in hindsight see that our proofs differ significantly and even the precise results proved are different."
Buckmaster remained unconvinced, writing on Mastodon that OpenAI's own explanation amounted to "openly admitting they used training data from a period after we found our result." OpenAI stated it began training the underlying system on August 28, claiming the model eclipsed internal benchmarks across complex mathematics.
Why It Matters
The dispute highlights a looming intellectual property tension at the intersection of AI tooling and cutting-edge academic research. As top researchers rely on cloud-hosted AI assistants to explore unreleased theories, the ambiguity surrounding how developer telemetry and de-identified session logs feed future foundation models creates severe trust issues between tech giants and the scientific community.
Reporting based on coverage from AI | The Verge.




