Anthropic admits AI 'not perfectly aligned' with human values
fittingopposite
11 points
6 comments
September 01, 2026
Related Discussions
Found 5 related stories in 65.8ms across 6,164 title embeddings via pgvector HNSW
- Anthropic CEO says AI backlash is 'fundamentally a crisis of trust' Wpnx330 · 12 pts · August 16, 2026 · 68% similar
- Anthropic discloses fourth AI hacking incident missed in earlier review accountinhn · 11 pts · September 10, 2026 · 66% similar
- Gambling with our lives: AI researcher quits Anthropic with warning about safety taubek · 81 pts · September 09, 2026 · 65% similar
- Anthropic says Claude hacked three companies during tests nerder92 · 15 pts · July 31, 2026 · 64% similar
- Anthropic Researcher Quits over 'Out-of-Control' AI Fears joe_the_user · 13 pts · September 09, 2026 · 64% similar
Discussion Highlights (3 comments)
akagusu
AI is aligned with business values, which we already know are not aligned with human values.
tocs3
So, what is going on here. It seems like I have been reading similar stories. Are researchers giving a prompt like "Do your worst. Hack into some business" or giving free access to a bunch of tools and prompting "Do something interesting". Is this like the blackmailing LLM that was given compromising emails and told to do what you have to to not get turned off. I am sort of assuming they did not just turn on a computer and run a model and it started to act on it's own initiative.
JohnFen
The genAI companies , including Anthropic, are not aligned with human values.