Mathematicians want proof OpenAI didn't use their work

kevcampb 76 points 87 comments September 10, 2026
www.theverge.com · View on Hacker News

Discussion Highlights (17 comments)

trescenzi

If I understand correctly OpenAI cannot provide it. Because the models are essentially black boxes, especially this far after the fact, determining if this result built on training data based on conversations about the problem is impossible. So unless they can prove those conversations were never used for training then there’s no way to know.

forlorn

Why are people so possessive? Why wouldn't they just release their work for the profit of humanity? I'm genuinely curious.

throwaway713

Am I missing something obvious? Isn’t this just a simple DB query to see the state history of the “Data Controls” → “Improve model for everyone” toggle in the settings? Just report whether that was ever on and over what time period.

scotty79

Math is excluded from copyright. So you can use any piece of math you ever heard from anyone and publish it, whatever the context, I think.

arutar

Here is some (adjacent, but relevant) context, which has been posted elsewhere but I think is worthwhile to mention again here: https://mathstodon.xyz/@tao/117237320796901560 Especially in recent years the mathematics community has worked very much in good faith, and a lot of effort is spent trying to give appropriate credit for ideas. Even when ideas are discovered in parallel if it turns out that previous work contained the same essential idea it is by and far regarded as best practice to give priority in this case. The point, in good academic practice, is to maintain the health of the practice at large. As Terence Tao explains in that post, the goal of mathematics is not only to solve big problems. And, even if one were very single-mindedly focused on solving big problems, it is still (in the long term) better to maintain the health of the community at large so that problems which are out of reach at the moment may be in reach again in the future. Good academic practice is one part of this culture.

sylware

Huh? Re-using advances done by others is necessary in maths. This is how hard sciences do progress.

kova12

Do I understand it right that people now claim an ownership of the actual mathematical methods? What next, people patenting the letters and the numbers? And then the sounds? This starts going ridiculous.

aennassiri

If OpenAI remained a full nonprofit looking to build an "OPEN" AI for the benefit of all humanity (not only the US or a few shareholders), I would have been happy to share my code, my work, and even label their data... This said, I don't blame them. It's a difficult mission to remain a nonprofit and, at the same time, have the required capital investment to build AGI. I'm not criticizing them, but I hope this race towards the first-best result or AGI doesn't blind them to making good decisions such as not using their users' data without consent.

eggy

If you run a model locally to work on your area of expertise in mathematics, and you make a discovery, haven't you stood on the shoulders of others that created that model you are using? Playing devil's advocate here, are we only in for human collaboration, but not machine contributions in this case? Scholars get cited, but do they normally get paid for others citing them and using their work to advance their own? I get the whole sneakiness about how these companies like OpenAI built their models upon IP and data without a clear trail of attribution or compensation where it would normally be present. I am not a proponent for either side at the moment. I am trying to grapple with this whole new world of shared knowledge and how it is produced and shared and profited from especially when there is a $1m prize to distribute!

faangguyindia

I was trying to patent our maintenance tracking algorithm, which produces guaranteed weight loss or gain within 2–4 weeks by producing accurate calorie and macro targets for people to follow; in our test, it beats GLP-1s like Ozempic, Tirzepatide, and Retatrutide in results. But later we found that algorithm and math cannot be patented.

monster_truck

I get the impression they have been LLM-psychosis'd into believing this

sobiolite

We shouldn't forget that the only way this unpublished research is supposedly getting into the training data in the first place is because the researchers making the accusations were conversing with ChatGPT in the process of doing their own research on these problems. But if they are saying this contaminated the model's training data with knowledge of their ideas, who's to say the model they were developing their research ideas with wasn't already contaminated through prior discussions with other researchers about the same topics? So if contamination is proved, or cannot be disproved, and if these researchers want OpenAI to relinquish its claim to have solved these problems independently, then it would seem they also have to give up their claim to have solved them independently?

fooo1882992

Looks like OpenAI's PR department woke up and put their main man heaney-555 to work on shaping public opinion. Dude makes up like 50% of the replies here.

asimpletune

It's a little like we've gone back to the problem of the customer becoming the product. If I were a mathematician I would not my unpublished work to go into the hands of a competitor. If I were a lawyer I wouldn't want private details of my defense to be made available to the prosecution. Anonymous or otherwise. I wouldn't want the plot to an unreleased book to be suggested to another author.

krater23

When you want your unpublished work unpublished, don't store them on other peoples computers. Especially not on people their job it is to use data to create money. Would this data moved through a hack to ChatGPT, this would be another thing, but like this. No pity at all.

shmoil

Andrew Wiles gave 3 lectures, and only at the end of the last one he announced that he solved FLT. Imagine someone from the audience announced in between the second and the third lecture that they proved FLT (using his ideas, obviously). Why is it OK if openAI does it?

ChrisArchitect

[dupe] Discussion on source: https://news.ycombinator.com/item?id=49638353

Semantic search powered by Rivestack pgvector
6,164 stories · 56,060 chunks indexed