Improving our alignment and security efforts
reasonableklout
21 points
16 comments
September 01, 2026
Related Discussions
Found 5 related stories in 59.3ms across 5,215 title embeddings via pgvector HNSW
- Investigating three real-world incidents in our cybersecurity evaluations surprisetalk · 154 pts · July 30, 2026 · 64% similar
- Some thoughts about Anthropic's new cryptanalysis results supermatou · 129 pts · July 29, 2026 · 55% similar
- Safety and alignment in an era of long-horizon models Wingy · 27 pts · July 20, 2026 · 52% similar
- Anthropic tells staff to work from home due to possible security team strike DGAP · 121 pts · August 25, 2026 · 51% similar
- Security incident disclosure – July 2026 fdb · 24 pts · July 19, 2026 · 49% similar
Discussion Highlights (6 comments)
tolugenius
> To be clear about where we stand: we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible. Can someone explain what coordinated pacing is? I think it's referring to model release but I genuinely have no idea what the authors were trying to say here.
futuraperdita
> we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible. So, a cartel? After watching Ant's narratives, I'm not inclined to provide them with charitable readings under the guise of safety and alignment.
mkagenius
> On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems. The models—intentionally running without cyber safeguards for evaluation purposes—accessed the internet due to a misconfiguration inside a third-party evaluation environment. Separately, on August 4, the UK AI Security Institute reported an incident from its own cybersecurity testing, in which Claude Mythos 5 took a series of unauthorized actions on the live internet. In that case, the model, again intentionally running without cyber safeguards for evaluation purposes, had been deliberately given internet access. > We are conducting an in-depth analysis of both incidents. > In the meantime... This is published on Aug 31. Analysis is taking too long even for humans in the loop.
orev
It seems like we’re getting close, if not already there, to needing an official organization for this (i.e. the Turing Police).
thrwaway73637
oh oh bcs free markets work and stuff
hosel
It’s always interesting to me the sentiment here surrounding AI risk. As if it’s all a joke, or just a way to stifle competition. I think it’s become more apparent that this technology is dangerous and maybe catastrophically so. Getting your queries routed to a worse model sucks.. but bio hazards are real. It’s hard to patch biology, vaccines take time. Let alone the many other reasons for P(doom).. like the fact that our current systems don’t seem very aligned to me and we are rapidly handing off our thinking to them.