TECHNOLOGY

OpenAI releases 372 math results from an internal model, including a claimed proof of the Unique Games Conjecture, Scott Aaronson reports

In a blog post, computer scientist Scott Aaronson describes OpenAI's release of hundreds of AI-generated math results. Humans have barely begun to check them, and the claims have not been independently confirmed.

Whiteboard covered in handwritten complexity theory notation in a university office, with a laptop showing a proof file beside itTECHNOLOGY

Image: IntraGoals Media · Uploaded by IntraGoals — usage rights confirmed

Computer scientist Scott Aaronson says OpenAI has released 372 mathematical results produced by an internal AI model. In a blog post dated October 7, 2026, he calls it one of the biggest days in the history of mathematics.

According to Aaronson, the release includes a proof of the Unique Games Conjecture, or UGC. It was proposed by Subhash Khot. Aaronson says the UGC implies that many optimization problems are truly NP-hard (meaning extremely hard to solve efficiently), even if you only want an approximate answer slightly better than what a standard technique called semidefinite programming relaxation gives. His wife, complexity theorist Dana Moshkovitz, has worked toward proving it for her whole career.

Aaronson writes that OpenAI released the results on the recommendation of its advisory group, which he says includes Timothy Gowers, Edward Witten and other mathematicians. The group later put out its own statement. It says its advisory role should not be read as a judgment of the results' impact, or as an endorsement of how OpenAI obtained them.

Nobody is sure yet that the proofs hold up. Aaronson says some of the results come with a Lean certificate, which is a machine-checkable formal proof, though not all of them do. He also says no human appears to have understood most of the proofs so far.

He shares texts from Moshkovitz about the UGC paper. She calls it badly written and says it is hard to read without AI help. She says it invents an unusual new kind of code, with a recursive construction, and that its citations are often irrelevant or confusing. She also says the authors give direct NP-hardness proofs for main applications of the UGC, such as Max Cut, that bypass the conjecture itself. She told him she still suspects a proof using a more natural tool, the half-space code, might exist.

Aaronson lists other results he plans to follow. They include L=BPL, a question about whether randomness helps in small-memory computation, and integer multiplication in less than O(n log n) time. He also lists a positive answer to the Unitary Synthesis Problem, which he and Greg Kuperberg posed in 2007, and a proof that parity is not in QAC0. Others are a near fourth-power separation between randomized and quantum query complexity, an area law for 2D gapped Hamiltonians, and matrix multiplication in O(n^(9/4)) time. He adds that the list includes uncomputability of solving polynomial equations over the rational numbers.

He says the release also includes partial progress on the Riemann hypothesis, the Hodge Conjecture and the Birch-Swinnerton-Dyer Conjecture. Some famous problems are missing. P≠NP is not on the list, nor are P=BPP or NEXP⊄P/poly.

Aaronson also mentions a separate result from the day before. Virginia Williams and Josh Alman posted an arXiv preprint giving faster algorithms for the 3SUM and All-Pairs Shortest Paths problems. He says an Anthropic model supplied the key idea. According to him, Anthropic gave the two researchers the chance to write and announce a digested version, in exchange for compensation, instead of posting raw solutions.

He sees tradeoffs in both approaches. OpenAI's way sets off a race among humans to digest and explain messy AI proofs. Anthropic's way puts a private company in the position of choosing which mathematicians act as the public face of the AI's work.

On how the results were produced, Aaronson says it was apparently not a custom setup of 10,000 agents, like the one he says was used for a finite-time blowup of the Navier-Stokes equations. Instead, he says it was the latest internal OpenAI model, which might reach paying ChatGPT customers within a couple of months, depending on its safety board. He says the company used about three hours of GPT-Pro-level compute per problem on average, and tried roughly 8,000 problems. That works out to about 5% solved on a single attempt.

Aaronson says researchers at the Simons Institute in Berkeley and at UT Austin are rushing to read and explain the manuscripts. He expects skeptics to call the results "AI slop" or to say the problems were trivial. He rejects that view, though his post is an opinion piece and not a neutral assessment.

He adds two updates. First, cryptography is missing from OpenAI's list, which he now says has 376 papers. He says his sources tell him AI companies have begun quietly checking whether their newest internal models can break important cryptographic protocols. Second, he quotes the advisory group's statement in full. It calls the release "the beginning, not the completion" of the process of human understanding, and says mathematicians must still be free to pursue their own questions. It also calls for equitable access to powerful research tools and enough computing resources.

Sources and further readingShtetl-Optimized ↗
ABOUT THE DESK

Harsh DV

IntraGoals reports on important changes in technology and work. We check each story for clear writing, trusted sources and useful information before it is published.

KEEP READING

Latest from IntraGoals.

All latest stories ↗