OpenAI: "We use ... de-identified data to improve ChatGPT"

nycdatasci 14 points 6 comments September 12, 2026
twitter.com · View on Hacker News

Discussion Highlights (3 comments)

nycdatasci

Post from Mark Chen for those not on X: Two things to distinguish: Did any human or agent look at user data as part of the Navier Stokes effort? No. Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company.

mmooss

> de-identified The issue always is, was the de-identification effective? With a birtdate, gender, and zip code, ~85% of Americans can be uniquely identified.[0] Much data contains much more unique information than that; I imagine most data about you has identifiable fingerprints - where you go, what you bought at the grocery store, your medical conditions, movies you watch, music you listen to, entertainment choices, hobbies, etc. An LLM is the perfect tool to identify someone based on that data.

zaptheimpaler

Man weve lived through this playbook already with Facebook and Google and an entire industry already over the last decade, we can’t be this naively trusting of entities which ultimately maximize shareholder/owner value and nothing else. Meta already pirated all the books in the world to train the models, lied about it and got caught. We know the models all stole copyrighted information and somehow got away with it - where even if it is fair use to train on it, they pirated it in the first place. Sam Altman is already known to have made serious lies throughout his career. OpenAI apparently doesn’t even know what websites their own damn models are hacking. It’s insane to trust these same people at their word now. The whole company will not know, someone at a high level could easily steal the data. Apparently even this tweet is saying they use your de-identified data to train on even when you explicitly turn off those options in settings. It’s the same old tricks again.

Semantic search powered by Rivestack pgvector
6,278 stories · 57,251 chunks indexed