
OpenAI says an internal AI system has solved the Navier-Stokes Millennium Prize Problem, one of mathematics’ famous $1 million challenges. The proof may turn out to be a historic AI achievement. It is not yet an accepted Millennium Prize solution, and another question now sits beside the mathematics:
Where did the model’s ideas come from?
The question became unusually concrete because mathematicians working on closely related results had spent months using OpenAI’s own Codex products, including with unpublished drafts. OpenAI says its researchers and agents did not access those researchers’ specific private data. In its initial account, the company also said it could not rule out de-identified data derived from their product usage having helped improve its models.
There is no evidence proving that the researchers’ private work trained the model that produced OpenAI’s proof. One of the mathematicians raising the concern explicitly says he does not know whether that happened.
The concern still exposes a problem that will become harder to avoid as scientists use frontier AI as a research partner. A company can provide the tool used to develop unpublished ideas, train future models on permitted user content, then deploy those models as researchers in their own right.
If AI starts competing with its own expert users for discoveries, a correct proof will answer only one part of the question. Research credit also depends on who supplied the decisive idea and whether anyone can audit that path.
Key takeaways
OpenAI has published a serious candidate solution, including a 166-page mathematical proof and a Lean formalization, but the Clay Mathematics Institute has not awarded or recognized the $1 million prize.
OpenAI says no specific user data from mathematicians Tristan Buckmaster and Levent Alpöge was accessed during the effort. Its initial account also said it could not rule out de-identified data from their use of OpenAI products having improved its models.
Buckmaster says he and Alpöge had been putting drafts from their closely related research into Codex. He says he asked OpenAI about training on those sessions and did not initially receive an answer. He also says he has no evidence their data was actually used.
OpenAI’s consumer data rules make the concern technically plausible in general. Personal ChatGPT and Codex content may be used for training unless the user opts out, while business products and the API are excluded by default.
The evidence no longer supports the simple claim that AI can only retrieve solutions it has already seen. Controlled math tests and a separately human-verified open-problem result show stronger capability. What remains weak is the ability to audit the intellectual provenance of a closed model’s discoveries.
What the OpenAI Navier-Stokes solution actually claims
On September 8, 2026, OpenAI published a proposed solution to the Navier-Stokes existence and smoothness problem. The company says its internal system constructed a smooth, externally forced three-dimensional fluid flow that starts at rest and develops unbounded velocity in finite time while retaining bounded kinetic energy.
That external force is important because it changes how many readers will instinctively understand the result.
The popular version of the Navier-Stokes problem is often phrased as asking whether a smooth three-dimensional fluid can spontaneously develop a singularity. The official problem description written by Charles Fefferman permits four routes to a solution. OpenAI claims to establish alternatives C and D, both of which permit a smooth external force.
So the forced construction is part of the official Millennium Prize formulation. It is not a workaround invented after the system found an easier neighboring problem.
OpenAI’s 166-page proof constructs the relevant finite-time blowup and says it establishes alternatives C and D. The paper also places the construction within a longer line of work, including research by Diego Córdoba and Luis Martínez-Zoroa on singularity formation through amplification across scales.
Calling this a “$1 million solution” still gets ahead of the process. The Clay Mathematics Institute continues to list Navier-Stokes among its unsolved Millennium problems. Its prize rules require qualifying publication, a waiting period of at least two years, and general acceptance in the mathematics community before Clay will consider a proposed solution.
OpenAI itself says it does not intend to claim the prize.
For now, this is a major proposed solution with formal verification, not a $1 million check waiting at reception.
This was an industrial research process, not one chatbot prompt
“An AI solved Navier-Stokes” compresses a strange and enormous research operation into five words.
OpenAI says it began the effort on September 1 after hearing rumors that two Millennium Prize Problems had been resolved. The company later connected the rumor to work by NYU mathematician Tristan Buckmaster and Anthropic employee Levent Alpöge.
Its internal system then attacked multiple Millennium problems using groups of agents with cached internet access and code execution. According to OpenAI, the Navier-Stokes group eventually involved roughly 10,000 concurrent agents.
The system first made progress on an easier Euler problem. OpenAI then redirected resources toward Navier-Stokes and supplied agents with that Euler result. Codex was used to consolidate promising ideas between agent groups. Newer checkpoints of the still-training internal model were introduced during the effort.
OpenAI reports that Navier-Stokes alone consumed about 2.7 million agent messages and 130 billion output tokens, with the proposed resolution emerging after roughly 88 hours. GPT-6 Astra then spent another 17 hours on Lean formalization and verification.
That scale changes how the achievement should be described. “Autonomous AI discovery” can mean one model receiving one prompt and working alone. It can also mean humans choosing targets, moving resources, feeding intermediate results between systems, updating model checkpoints, coordinating thousands of agents, providing retrieval and code execution, and using another model to formalize the result.
Those are different experiments. Both may be scientifically interesting, but they answer different questions about what the model itself can do.
The Navier-Stokes effort is best understood as AI-led research infrastructure operating at industrial scale. That does not make the result less impressive. It does make the provenance trail more complicated, because many systems, prompts, intermediate results, human choices, model versions, and retrieved documents can influence the final path.
Another team was working nearby with Codex
Buckmaster and Alpöge were simultaneously developing related singularity results.
In a public statement describing their work and their use of Codex, Buckmaster says their project built on the program developed by Córdoba and Martínez-Zoroa. He and Alpöge used several LLMs during the project, including Anthropic’s Claude and OpenAI Codex, especially GPT-5.6 Sol, with Astra later used for writing and auditing.
Buckmaster says the pair obtained smooth-forcing blowup results for Boussinesq and Euler on August 15 and verified the Euler result in Lean on August 22.
Then came the coincidence that set off alarms.
OpenAI’s result also took a smooth-forcing route. Buckmaster regarded that direction as surprising because it was closely related to the research program his group had quietly pursued. More important for the data question, he says the team’s drafts had been placed into Codex throughout the project.
During discussions with OpenAI on September 6, Buckmaster says he asked whether the internal model had either accessed or been trained on those Codex sessions. According to his account, he was told that the model had not looked up user data. When he asked specifically about training, he says he did not initially receive an answer.
Buckmaster is equally explicit about the limit of his evidence. He says he had not seen OpenAI’s proof, did not know how its model obtained the result, and did not know whether his team’s data was used.
He says he was not accusing OpenAI of having done so.
That limit is central to the story. Similarity in research direction is not evidence of data leakage by itself.
There is also a strong reason two groups could converge without private-data access. Both openly credit the same published mathematical lineage. Buckmaster says Córdoba and Martínez-Zoroa supplied the basic ideas behind his program. OpenAI’s proof cites their work and describes its own use of dynamical amplification inspired by that research.
Researchers who start from the same literature can arrive at related routes. Mathematics is full of simultaneous discoveries for exactly that reason.
The uncomfortable part is that Codex sat inside one team’s private workflow while the company running Codex was also building a system capable of competing on the same class of problems.
OpenAI denied direct access, but its initial caveat left a provenance gap
OpenAI says its researchers and agents did not see Buckmaster and Alpöge’s unpublished work before its public release and that no specific user data was accessed to solve Navier-Stokes.
Its initial public account went further in a less reassuring direction. OpenAI said it could not rule out that de-identified data derived from the mathematicians’ use of OpenAI products had helped improve its models.
That statement was never evidence that their work was used. It was an admission that OpenAI could not provide the stronger guarantee readers might reasonably want in a case involving unpublished mathematics and a historic research claim.
The difference is practical. “The agents did not open these users’ private sessions” addresses direct retrieval. It does not, on its own, answer whether permitted training data from earlier product use could have shaped a later model.
For ordinary consumer use, that gap can feel abstract. For a mathematician putting unpublished proof sketches into a hosted coding agent, it can become a priority and attribution problem.
A model does not need to reproduce a private draft verbatim for the draft to matter.
A training process could, in principle, absorb a technique, a promising route, a counterexample shape, a correction, or a failed approach that steers later reasoning.
The evidence does not establish that any of those things happened here. The point is that a closed training pipeline makes the question hard to audit from outside.
The control lever is permission to train on expert users
OpenAI’s current data-use documentation says content submitted through individual services such as ChatGPT and Codex may be used to train its models. Users can opt out.
Codex adds another control surface. OpenAI says full environments have separate training controls in Codex Settings, and changing the normal ChatGPT setting or using the privacy portal does not alter those full-environment controls.
OpenAI says ChatGPT Business, Enterprise, and its API work differently. Inputs and outputs from those business services are not used for model training by default unless the customer opts in.
Its Terms of Use say users retain ownership rights in their inputs and own their outputs while allowing OpenAI to use content to provide, maintain, develop, and improve its services, subject to the available controls.
Those rules can coexist without contradiction. Ownership of a document and permission to process that document are separate questions.
For research, though, contractual permission does not settle scholarly credit. A scientist can agree to product terms without intending to donate an unpublished idea to a future research competitor. The legal question may be whether the processing was allowed. The academic question is whether an idea that materially contributed to a later result deserves attribution.
De-identification does not automatically settle that question either. OpenAI’s privacy policy says it may aggregate or de-identify personal data and use that information to improve features and conduct research. Removing a person’s name, account ID, or other identifying information from a mathematical proof sketch can leave the mathematics intact.
Identity and intellectual content are different objects. A theorem does not stop being informative because its author’s name has been deleted.
Again, none of these policies prove Buckmaster and Alpöge’s work entered the model. They explain why researchers should understand the training controls before placing unpublished work into a hosted AI system.
The ethics change when the tool provider can compete with its users
AI companies want scientists, engineers, and programmers to bring hard problems into their products. Hard work creates valuable interactions. Users get capable assistants, while model developers learn where the systems fail and where they improve.
That arrangement becomes harder to evaluate when the provider also deploys its own models as researchers.
Imagine a hosted research tool learning from thousands of mathematicians’ private attempts, partial proofs, failed constructions, intuitions, literature searches, and corrections. No single conversation needs to contain a complete solution. The value may sit in fragments scattered across many sessions.
Later, a model from the same provider solves a related problem.
Academic credit has established machinery for papers, citations, collaborators, prior art, private communication, and dated drafts. It has much less machinery for a model whose useful mathematical knowledge may have been distilled from millions of interactions without retaining an inspectable path from output back to contributors.
That is the real control problem. The provider owns the service, controls the data pipeline, chooses the training process, decides what records are retained, develops the research model, and can publish what that model produces. The expert user cannot independently inspect any of those layers.
Contractual permission to train is therefore only one part of the story. Research provenance also needs records that can support or falsify a claim of independence.
If the decisive insight came from ordinary published literature, attribution can work in familiar ways. If it came from a user’s private interaction and was absorbed into training, conventional citation systems have no reliable mechanism for recovering that path after the fact.
OpenAI’s result is not independent of human mathematics, and it does not need to be
There is another problem with the word independent.
OpenAI’s agents had access to a cached internet. Its proof cites decades of human mathematics. The successful research process used the agents’ Euler result as a stepping stone and moved insights between agent groups.
None of that disqualifies the work from being a discovery.
Human mathematicians read papers, reuse lemmas, learn techniques, ask colleagues for help, attack easier variants, combine ideas from different fields, and build on generations of prior results. Requiring an AI to reinvent every prerequisite from first principles would set a standard no human researcher meets.
The useful test is whether the target solution, or a decisive unpublished insight needed to reach it, was already available to the system in a way that collapses the claim of discovery.
That is why retrieval history matters. A model that finds an obscure existing proof has done something useful, but it has not solved the open problem. A model that constructs a new proof from published ingredients has done something much closer to research, even though the ingredients are human.
Recent AI math results make the second possibility increasingly hard to dismiss.
We already saw what mere retrieval looks like
There is good reason to be skeptical because OpenAI has previously overreached on mathematical novelty.
In October 2025, an OpenAI executive said GPT-5 had found solutions to 10 previously unsolved Erdős problems. Mathematician Thomas Bloom objected that the characterization was wrong. As TechCrunch reported, GPT-5 had located existing literature containing solutions that Bloom’s database had not yet recorded.
That was impressive literature search.
It was not independent mathematical discovery.
The episode gives us a useful baseline. If the answer exists somewhere in the accessible literature and the model finds it, “AI solved an open problem” is the wrong headline, however difficult the search was.
The harder test is whether a model can construct a valid answer when the target proof is unavailable.
There are now better examples of that.
First Proof makes the retrieval explanation harder
The First Proof project released 10 research-level questions while initially withholding the authors’ solutions. Those solutions were later released together with the keys to previously encrypted versions. The organizers explicitly said solutions completed before the official answers became public would carry the greatest credibility.
That setup attacks the simplest retrieval explanation. If the target proof has not been released, a system cannot merely find the authors’ answer online.
OpenAI ran an internal model on all 10 questions. In its February report, OpenAI said expert feedback gave at least five attempts a high probability of being correct.
The experiment also produced a useful failure. OpenAI initially thought another proof was probably correct, then acknowledged that further review showed it was wrong.
That correction matters because research-grade mathematics cannot rely on the model’s confidence or the lab’s first impression. Proofs still need checking.
OpenAI also conceded that its evaluation was not perfectly controlled. Humans sometimes encouraged promising strategies, selected among attempts, and requested clarifications. That makes “fully autonomous” a stronger description than the setup supports.
Even with those caveats, correct solutions produced before the official answers became public are evidence for mathematical construction beyond simple answer retrieval.
They do not show intellectual isolation from everything the model learned in training. No model trained on human mathematics could meet that standard. The relevant question is whether the system built a new solution from prior knowledge rather than recovering the target answer.
A separate open problem has survived human verification
There is stronger evidence from another 2026 result.
In May, OpenAI announced that an internal model had disproved a longstanding Erdős conjecture concerning unit distances in the plane.
A group of prominent mathematicians including Noga Alon, Tim Gowers, Will Sawin, Arul Shankar, Jacob Tsimerman, and others subsequently published a human-verified treatment of the model-generated counterexample.
Their paper traces the argument to known mathematical ingredients while describing the counterexample as OpenAI-generated. The result is useful evidence against the strongest claim that frontier models can only repeat complete answers already present in the literature.
A model can use known mathematics and still create a novel combination. Humans do that every day. The real research question is whether the combination resolves something that was genuinely open and whether the argument survives independent checking.
The unit-distance result cleared a much stronger external check than a lab simply announcing that its own model had succeeded.
That makes the provenance problem more urgent. As models become capable of real mathematical construction, knowing what they were exposed to before the discovery becomes more important.
Correctness and provenance need separate audits
Mathematics has an unusually clean mechanism for checking an answer. A proof is valid or it contains a flaw. Formal verification can remove a great deal of ambiguity from that question.
Formal verification cannot tell you where the idea came from.
A Lean checker can establish that a long chain of steps follows correctly. It cannot establish whether an unpublished human sketch influenced a model months earlier through training, whether a researcher supplied the key strategy in a prompt, or whether an agent retrieved a decisive source during the run.
Future claims of autonomous AI discovery therefore need two different audits.
A correctness audit asks whether the proof works. Publish it, formalize what can be formalized, and let independent experts attack the argument.
A provenance audit asks what the system knew and how the research run was constructed. That means preserving records of when the problem became available to the model, what the agents retrieved, which prompts and intermediate results were supplied, what human interventions occurred, which model checkpoints were used, and what training-data boundaries applied to those checkpoints.
Those records do not need to reveal every proprietary training example to be useful. Even coarse but auditable boundaries would be better than forcing outside researchers to infer independence from a company’s assurances after a dispute begins.
The Navier-Stokes case shows why the two audits cannot substitute for each other. A formally verified proof could still have a messy provenance history. A perfectly documented provenance trail could still end in a wrong proof.
Scientific credit needs both questions answered separately.
Researchers should treat hosted AI as part of the publication threat model
The immediate lesson is practical for anyone using AI on unpublished work.
OpenAI’s policies provide more control than many users realize. On personal ChatGPT and Codex accounts, researchers should check whether model improvement is enabled and separately inspect Codex’s full-environment settings. For work requiring stronger default boundaries, OpenAI says its business products and API do not train on customer inputs and outputs by default.
Popular AI’s AI privacy and security guide explains how to separate ordinary prompts from intellectual property and sensitive professional material. That separation is useful even when a provider has strong policies, because the easiest data leak to fix is the one that never enters the wrong system.
For genuinely confidential discoveries, the safest architecture is still one where the crucial material never enters a training-eligible hosted environment. A local AI setup can handle supporting tasks without placing the same unpublished material into another company’s hosted training pipeline, though local models may not match every frontier system on difficult research.
Researchers using hosted frontier models should also preserve dated drafts, prompts, model outputs, Lean files, Git history, emails, and local working notes. Those records can become evidence of priority when human and machine contributions overlap.
The broader AI autonomy question is also a control question about who holds the records needed to establish what a model knew and when it knew it. A research lab that controls the model, the service, and the logs has an evidentiary advantage over the individual researcher using the product.
That does not mean researchers should stop using hosted AI. It means unpublished work deserves the same deliberate handling as source code, patentable ideas, embargoed results, or confidential client material. Training controls are part of the research workflow now.
More on AI privacy:
The OpenAI Navier-Stokes solution needs a provenance standard to match the proof
OpenAI appears to have produced something much more serious than another chatbot claiming it solved a famous theorem. There is a detailed proof, a Lean formalization, an enormous disclosed agent effort, and enough evidence from other mathematical experiments to reject the idea that frontier AI can only parrot complete solutions already available online.
Whether the Navier-Stokes proof survives the mathematics community’s scrutiny is still unresolved. The $1 million prize has not been won.
At the time of the initial announcement, the specific data question also remained unresolved. There was no evidence that OpenAI took Buckmaster and Alpöge’s private Codex work and fed it directly to the agents. Buckmaster did not claim to know that happened. OpenAI denied specific user-data access.
The concern arose because OpenAI’s initial account could not provide the strongest possible assurance about de-identified product data and because the company simultaneously operated the research tool, controlled the training pipeline, and built the competing research system.
That combination should raise the standard for future claims of autonomous discovery.
A correct proof demonstrates mathematical capability. A credible claim of independent discovery requires a separate evidentiary trail showing that the system did not quietly inherit the decisive unpublished idea from the people whose work it is now competing with.
AI research systems are getting good enough that this will not remain a theoretical problem. We may soon be able to verify exactly what a model proved while still lacking a reliable way to reconstruct who supplied the idea that made the proof possible.
The fix is not to demand that AI forget human mathematics, merely because it has borrowed conclusions from human predecessors. It is to make the research path auditable enough that published knowledge, private user work, retrieved material, human intervention, and model-generated reasoning can be separated after the fact.
If labs want credit for autonomous discovery, they should be prepared to show more than a correct answer. They should be able to show the chain of custody for the ideas that produced it.
Explore more from Popular AI:
Start here | Local AI | Builds & gear | Autonomy & policy | Fixes & guides | Popular AI podcast














