OpenAI: "We use ... de-identified data to improve ChatGPT"
nycdatasci
14 points
6 comments
September 12, 2026
Related Discussions
Found 5 related stories in 137.3ms across 6,278 title embeddings via pgvector HNSW
- How Organizations Use AI: Evidence from ChatGPT [pdf] malshe · 91 pts · August 13, 2026 · 67% similar
- A milestone in expanding access to AI tosh · 12 pts · August 31, 2026 · 66% similar
- ChatGPT Work Tiberium · 332 pts · July 09, 2026 · 64% similar
- OpenAI restores 5-hour Codex and Work limits for ChatGPT Plus users MC995 · 112 pts · August 25, 2026 · 64% similar
- ChatGPT Images 2.5 vertigoruntime · 310 pts · September 08, 2026 · 63% similar
Discussion Highlights (3 comments)
nycdatasci
Post from Mark Chen for those not on X: Two things to distinguish: Did any human or agent look at user data as part of the Navier Stokes effort? No. Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company.
mmooss
> de-identified The issue always is, was the de-identification effective? With a birtdate, gender, and zip code, ~85% of Americans can be uniquely identified.[0] Much data contains much more unique information than that; I imagine most data about you has identifiable fingerprints - where you go, what you bought at the grocery store, your medical conditions, movies you watch, music you listen to, entertainment choices, hobbies, etc. An LLM is the perfect tool to identify someone based on that data.
zaptheimpaler
Man weve lived through this playbook already with Facebook and Google and an entire industry already over the last decade, we can’t be this naively trusting of entities which ultimately maximize shareholder/owner value and nothing else. Meta already pirated all the books in the world to train the models, lied about it and got caught. We know the models all stole copyrighted information and somehow got away with it - where even if it is fair use to train on it, they pirated it in the first place. Sam Altman is already known to have made serious lies throughout his career. OpenAI apparently doesn’t even know what websites their own damn models are hacking. It’s insane to trust these same people at their word now. The whole company will not know, someone at a high level could easily steal the data. Apparently even this tweet is saying they use your de-identified data to train on even when you explicitly turn off those options in settings. It’s the same old tricks again.