OpenAI’s Astra Model Just Cracked 10 Math Problems That Sat Unsolved for Decades

OpenAI says its next model, Astra, produced ten proofs for math problems that had been open for a decade or more. Total compute cost: about $2,000.

OpenAI’s Astra Model Just Cracked 10 Math Problems That Sat Unsolved for Decades

OpenAI says its Astra model just solved ten problems in mathematics and theoretical computer science that had been open for at least a decade, most of them much longer. The whole run cost roughly $2,000 in compute at Sol API rates.

The company published the results on its research page on August 1. Each proof ships with a machine-checkable Lean 4 certificate on GitHub, plus a model-written narration of how it got there.

The full manuscript runs 249 pages. This is the first time OpenAI has publicly described what the Astra model can do. It calls the model its next major release, and it hasn’t said when anyone outside the company gets to use it.

What the Astra model actually solved

The list isn’t a set of toy problems. It includes the first explicit construction of a non-sofic group, a question in group theory that has stood since Gromov introduced the concept in 1999. Nobody had resolved it in 27 years.

The Astra model also disproved Connes’s rigidity conjecture, a long-standing problem about von Neumann algebras. It settled Ehrhart’s volume conjecture in every dimension. It resolved three open questions from Paul Erdős’s famous problem list, including problem 183 on multicolor Ramsey numbers.

Some results have real practical edges. The closest vector problem result tightens what’s known about a lattice problem that post-quantum cryptography leans on. The new arithmetic circuit lower bounds for the permanent push complexity theory forward.

The sphere packing bound is the first improvement to the general upper bound since 1978. OpenAI’s head of math research, Sebastien Bubeck, confirmed the news on X, posting that non-sofic groups exist and calling the batch of results beautiful.

Why the $2,000 number matters

The cost is the part that should stop you. Ten research-level results, some of them famous open problems, for the price of a used laptop. OpenAI says the token cost at Sol API rates for finding the solutions was around $2,000. That number frames everything.

This isn’t AlphaProof solving an IMO problem with a bespoke pipeline. This is the next general frontier model doing research-level math as a side effect of being good at reasoning.

The same model family already disproved the Erdős unit distance conjecture in May, an 80-year-old problem in discrete geometry. Tim Gowers said then that he would recommend that proof for publication in Annals of Mathematics without hesitation.

Thomas Bloom, who runs the Erdős problems site, called these ten results bigger news than the unit distance counterexample.

The math community has already started building on the May result. The footnote on OpenAI’s page lists follow-up papers on the sum-product conjecture, the Elekes-Ronyai problem, and the Minkowski grid, several of them uploaded to arXiv within weeks.

The verification question is the real story

The reason this matters more than a flashy demo is that the proofs are checkable. Lean certificates let anyone with the Lean compiler verify each argument without trusting OpenAI or its model.

That directly answers the objection the math community raised in the Leiden Declaration, the June statement warning that AI companies were bypassing peer review and threatening the integrity of proof and attribution.

OpenAI is aware of the tension. Its own post says claiming human authorship for an AI-generated proof would misrepresent both the system’s contribution and genuine human intellectual work.

The company says it prepared the manuscripts and formalized the proofs in Lean, and it takes responsibility for correctness. The mathematical arguments themselves were generated by the system.

That’s the honest framing, and it’s also the open wound. A blog post with machine-checked proofs isn’t the same as the peer-reviewed record, and the Leiden Declaration was written precisely because the community doesn’t want companies setting the rules for what counts as a result.

Whether journals accept Astra’s proofs will be a bigger test than the proofs themselves.

What this says about the AI race

Every major lab is chasing the same thing: models that do real work for you over long horizons. OpenAI previewed the Astra model in Washington recently, touting its ability to run multiple agents together for extended periods on hard problems.

The math results are the same capability pointed at research instead of code.

I’ve watched OpenAI’s pricing moves closely this week, since the GPT-5.6 price cuts reset expectations for what frontier intelligence should cost. This announcement is the other side of that coin.

Cheap inference isn’t just about making agents affordable. It’s about making discovery affordable.

Caveats first: ten proofs is a sample, not a guarantee. The results haven’t been through traditional peer review, and the safety questions around long-horizon agents don’t disappear because the model is doing math.

But the direction is unmistakable: the cost of original research just collapsed, and the verification tools exist to keep the field honest.

The Astra model is coming. When it ships, the interesting question won’t be what it scores on a benchmark. It’ll be what the math community decides to do with proofs it can check but did not create.

Tony Simons

Reviewed & Written By

Tony Simons

Independent tech reviewer and creator of Tony Reviews Things. 14 years of hands-on testing, software auditing, and workflow automation. I test the gear so you don't waste your money on junk.

Submit a Take

Your email address will not be published. Required fields are marked *