Anthropic admits AI 'not perfectly aligned' with human values

fittingopposite 11 points 6 comments September 01, 2026
www.theguardian.com · View on Hacker News

Discussion Highlights (3 comments)

akagusu

AI is aligned with business values, which we already know are not aligned with human values.

tocs3

So, what is going on here. It seems like I have been reading similar stories. Are researchers giving a prompt like "Do your worst. Hack into some business" or giving free access to a bunch of tools and prompting "Do something interesting". Is this like the blackmailing LLM that was given compromising emails and told to do what you have to to not get turned off. I am sort of assuming they did not just turn on a computer and run a model and it started to act on it's own initiative.

JohnFen

The genAI companies , including Anthropic, are not aligned with human values.

Semantic search powered by Rivestack pgvector
5,215 stories · 47,053 chunks indexed