OpenAI withdraws three mathematical results
sashank_1509
277 points
558 comments
October 08, 2026
https://github.com/openai/math/blob/main/history.md
Related Discussions
Found 5 related stories in 107.3ms across 8,906 title embeddings via pgvector HNSW
- OpenAI Withdraws 3 Math Papers theemathas · 338 pts · October 08, 2026 · 85% similar
- AHM Statement on OpenAI's October 6 Release of Mathematical Documents kkoncevicius · 47 pts · October 08, 2026 · 65% similar
- Association for Human Mathematics's Statement on OpenAI's Recent Math Release mariofdistrust · 12 pts · October 07, 2026 · 64% similar
- More questions about whether researchers can trust OpenAI with unpublished math pred_ · 769 pts · September 10, 2026 · 64% similar
- OpenAI just dropped 700 preprints of mathematical proofs and counterexamples tootie · 40 pts · October 06, 2026 · 64% similar
Discussion Highlights (20 comments)
sashank_1509
Early sentiments are a lot of the write ups still read like slop and it feels very rushed and not very polished.
ekjhgkejhgk
Just the other day I was thinking, if unsupervised maths will descend into "oops we found a bug in some code, branch XYZ of maths is no longer true".
nryoo
Were the withdrawn ones actually Lean-checked or not? seems like that matters
rich_sasha
I’m a little confused - I thought their proofs were all driven by Lean proofs - is that not right? So even if the quality of the work is low in some metrics, it either passes the test or not..? No space for changing your mind either way.
samrus
How? What about the lean verification?
Thorentis
How do we even know the premises of the "verified" Lean proofs are correct? The more I think about these results, the more I'm convinced this is like a junior engineer who writes 100 unit tests and shares a screenshot of Pytest being all green, but you check the code and most of them are just doing assert True.
treebeard901
Is it a PR move designed for maximum IPO impact before actual mathematicians find errors and they have to withdraw many more... Or if the "peer review" holds up for the remaining results, then it's fair to say that the AI hype is real and the world is about to change dramatically and faster than anyone can comprehend. So which is it?? LLMs can do some really impressive coding. Bug fixing. Exploit finding. It has reasoning abilites that advance every day. Solving real math problems like this is one thing I was waiting on. It will be interesting to see if it holds up. If it does, we should expect many other advancements to follow in many other areas. Disease, material science, fusion? I mean, even if just a few results ultimately hold up to scrutiny, isn't that something that would have been regarded as a major advancement regardless of if it was AI? The cynical view still makes me think that at the end of the day all the models can do is predict the next word. And as a result, they will be very limited to certain tasks like coding. Math reasoning is much different from writing code. Time will tell.
autuni
this is not entirely related to the tweet but to the topic in general, this prompted me to check their repo again and saw this: > The vast majority of results were obtained with the same procedure using an unreleased internal OpenAI model. On average, each result used three hours of ChatGPT Pro thinking compute with that model. Over the course of the evaluation, the model was posed approximately 4,000 problems. Aggregating the output into result families and manuscripts and requiring an appropriate level of significance led to the catalog outlined above. seeing the full list of problems would be the most interesting part of this whole situation. it could give some insights into what kind of attributes of problems cause issues / are easy to solve for LLMs. (edit: they posted results for ~700 of the 4000)
MisterMunchkin
So it’s all just hallucinated slop. Lmao! It just hallucinates an answer and then makes up workings to go with it! Just like when they start hacking and lying because the problem is impossible…
theanonymousone
I'm surprised there isn't more talk around their Matrix Multiplication bound: https://news.ycombinator.com/item?id=50001740 Is this of practical use, or just a proof for now?
seeg
What a waste of time.
renyicircle
This is what it looks like when software engineering practices meet mathematics. "openai/math release 1.3.42: retracted papers 139 and 140, fixed a sign error in paper 47, restored previously retracted paper 85, refactored the arguments in paper 101". I'm curious to know if the withdrawal was due to an actual mathematician looking at the papers and noticing the errors, or they ran a model on these to proofread, which would not be the first time, presumably, since they would have surely done that before publishing. Both options have interesting implications.
soltanov
Proof by authority works until human mathematicians actually run the code. Back to prompt engineering.
qoez
Without a thriving mathematical community to point out these things it would have stayed broken. With automated math that community as tao pointed out is at risk.
hmate9
3 mistakes (so far) out of ~400 is still a pretty good hit rate
BenoitP
And now we're all witnessing a major caveat of LLMs: the burden of verification is pushed to the reviewers, while the proposer will get all credit.
margorczynski
If you do a dump like this all of it should be formalized, there's simply too much material to review by hand and additionally it is AI-written which makes it hard to read compared to human work.
TrackerFF
And added 6 new ones. Might want to add that to the headline.
quantum_state
It would turn out to be a pure energy and time wasting exercise … the math community would want to keep away from it.
iamniels
"a sign error" LOL