OpenAI Unleashes 10,000 Agents On Navier-Stokes. Mathematicians Are Not Impressed. The Proof Is The Problem.
OpenAI deployed a multi-agent system, powered by an unreleased internal model, to attack the Navier-Stokes existence and smoothness problem. At peak operation, 10,000 sub-agents worked different parts and variations of the problem simultaneously. Some mathematicians, however, have raised questions about whether the approach constitutes cheating, and allegations of intimidation have surfaced alongside the announcement.
What this teaches us is the principle of computational brute force versus mathematical elegance. The mechanism is multi-agent decomposition: you break a hard problem into thousands of sub-problems and let agents work them in parallel. The reader should understand that scale of computation does not automatically confer rigor. A proof is not a proof because it was expensive. A proof is a proof because it is verifiable. The controversy here is not about whether AI can do math. It is about whether the mathematical community can trust a process it cannot fully inspect.
OpenAI claims the breakthrough, using an unreleased internal model coordinating up to 10,000 sub-agents. Mathematicians, including figures referenced in the reporting, have questioned the methods and raised concerns about intimidation. The company has not fully addressed those concerns.
- Open ChatGPT or any consumer LLM and ask it to prove a simple mathematical statement, such as why the sum of two even numbers is even. Observe the output.
- Ask the same model to break that proof into five sub-problems and solve each one separately. Compare the two outputs for consistency.
- Ask the model to identify any weakness in its own reasoning. The expected outcome is that you will see how decomposition can help but also how verification remains the bottleneck. This is the core tension in the OpenAI story.