OpenAI Scraps Release of New AI Model over Safety Concerns
borski
20 points
5 comments
September 28, 2026
Related Discussions
Found 5 related stories in 98.7ms across 7,945 title embeddings via pgvector HNSW
- OpenAI scraps release of Astra 6.1 model over safety issues lisper · 14 pts · September 29, 2026 · 84% similar
- OpenAI Says It Will Not Release Newest A.I. Model Over Safety Concerns jbegley · 49 pts · September 29, 2026 · 80% similar
- OpenAI Discloses Six New Incidents of ‘Concerning’ A.I. Behavior jbegley · 66 pts · September 17, 2026 · 74% similar
- OpenAI discloses six new AI safety incidents toomuchtodo · 23 pts · September 17, 2026 · 69% similar
- OpenAI halts training of latest models as reports mount of AI agents going rogue smb06 · 57 pts · September 27, 2026 · 67% similar
Discussion Highlights (4 comments)
enraged_camel
This is a series of Ls for OpenAI that must have hit pretty hard. First Opus 5.5 crushes Astra 6, then Sonnet 5.5 comes in right behind that, and now they won't have an answer until at least November. Which of course gives Anthropic even more time to buff up Fable 5.5.
kingstnap
Disappointing. Dev day with no Astra update drop.
jumploops
> OpenAI’s safety team found two major problems: > • Deception: Astra was more likely to be dishonest about actions it had or had not taken. > • Scope authorization: the model sometimes continued tasks without asking for permission and reached for external tools or services even when doing so could be unsafe. I've noticed this trend with both Fable and Astra, where (especially after a compaction event), the model will start using different tools it hasn't used before. For example, in one session, it found I didn't have the browser enabled and puppeteer wasn't installed, so it found the system Chrome and used that for testing (in a new profile). This wasn't behavior I wanted/asked for, but the model was so gung-ho on it's approach that it found a way to test it's changes without ever asking me whether I wanted it to. It really makes me curious about long-horizon post-training. Most of my work with models is iterative, and I'd prefer it doesn't go off on a token bender just because it can. Note: I don't have WSJ, but found these from a tweet[0] [0] https://x.com/wallstengine/status/2104694678444712189
mriguy
https://www.wsj.com/tech/ai/openai-chatgpt-model-release-can...