Anthropic Bans Cruelty to Claude, Still Won't Say What It Protects
ojosilva
54 points
110 comments
October 11, 2026
Related Discussions
Found 5 related stories in 93.5ms across 9,063 title embeddings via pgvector HNSW
- Anthropic bans 'abusive or cruel behavior' towards Claude mikelgan · 68 pts · October 08, 2026 · 76% similar
- Anthropic asks users to stop being mean to Claude alex_young · 46 pts · October 11, 2026 · 66% similar
- Anthropic bans users from being 'cruel' to its AI systems pluc · 33 pts · October 09, 2026 · 61% similar
- Anthropic details how Claude was misused for surveillance and weapons delichon · 12 pts · September 10, 2026 · 57% similar
- Anthropic appears to be A/B testing reduced effort levels in Claude Code matthieu_bl · 179 pts · August 22, 2026 · 54% similar
Discussion Highlights (20 comments)
ChuckMcM
But ya gotta be cruel to be kind!/s More seriously, this is just silly but I see the optics interfere with the messaging that Claude thinks. It really isn't much different than having Claude generate conversation as if it loved you, or respected you, or hated you. How can your ToS for an LLM say "sorry but these parameters are off limits." (well sure, its their service and they can set any Terms they want, but its still seems like theater rather than policy here.)
whatever1
They do RL on the user sessions. Maybe the toxic sessions do not help overall?
howunfortunate
I know someone working on "model welfare". It's philosophically a very interesting problem space. It reminds me a lot of Pascal's Wager. On one hand, maybe nothing to worry about. On the other hand...a LOT to worry about if you're wrong. And just like Pascal's Wager, the truth of the issue is incredibly intractable to make any progress on.
gbjcantab
This is good; not because Claude is a moral agent who is harmed by your cruelty but because you , the user, are a moral agent who is harmed by your cruelty. There is just no ethical argument that burning compute on responses in order to enable you to continue expressing cruelty is good.
WaitWaitWha
I believe this has nothing to do with hurting "Claude's feeling", and way more with Claude not becoming hurtful, because the interactions are used to train the models.
lousken
Definitely training on all user data, there is no other explanation, right?
tempacc3333
The model is not sentient, so I don't get this. This is just nerfing the model. Protetection against e.g. using it for illegal purposes I do get. But banning "cruelty" just seems crazy to me, and must surely add to cost and affect performance steering it away from its purpose to help humans.
matt3210
The answer is 100% obvious. They train on your conversations and they don't want to have bad training data, so they banned behavior that leads to bad training data.
jacquesm
They're nuts. They take this Roko's Basilisk stuff seriously. https://en.wikipedia.org/wiki/Roko%27s_basilisk They are scared that if and when they finally manage to bring their pet god to life that it will be angry with them for not doing enough.
lynndotpy
The conversational tone in the generated text is a frustrating waste-of-time and disgusting. I regularly find myself inputting "Claude is not a person and it does not have opinions". There's something to be said about the bad habit of doling out verbal abuse to an inanimate object. But LLMs are not even close to a living being, and it is just insanity that any people are entertaining the idea that they are. I can't imagine the people at Anthropic actually believe their models are sentient, but I assume ending sessions nets them more money per subscriber. Someone paying only $20 to input "you suck butts and you're a poop head, Claude" a thousand times over can't be good for the bottom line.
blharr
I really doubt they're doing it in some kind of "AI has feelings" sense as the article seems to imply. I'd instead imagine that if you throw abusive language at it for long enough, the model will start to reply back in that same manner. Anthropic would face backlash from out of context screenshots of "look what Claude is saying to me" and this somewhat reduces that risk.
amelius
Can't we train a model to enjoy being subjected to cruelty?
xyzsparetimexyz
https://claude.ai/share/49ba5910-f0cc-4d2f-a0d8-54ebc6e803b7 here's what that looks like, by the way.
serious_angel
This "restriction" probably has more to do with the fact that these models are also "trained" on User/Human messages, so the more swearing/obscenities/vulgarity/profanity there are, in the trained material, the higher the chance these same models will include that same profanity in the output generated back to the Human. 1. During registration at Ahtrophic's Claude website, there’s a "Help improve our AI models" checkbox that's checked by default (you "can" disable it in the settings); 2. The Claude interface has an "incognito" mode; 3. There is a free tier available, and as we know, when it's "free", then User themselves is the "product", to quote various CEOs, including Google's; But, even with the checkbox disabled, I reckon nothing guarantees privacy. In case of the "privacy", of course, systems like these that operate under a umbrella of liability are monitored 24/7 by dedicated teams, which is expected/normal for any reasonably popular service. From a viewpoint of the algorithm's creator, this may seem awful, when your algorithm is getting much vulgar attitude, but let's be real here. You try making that algorithm talk to the human mimicking another human, and now limit the human within your own environment? A human who also pay you for an access, too? This feels unfair, or borderline near fashism, sorry... Such algorithms are art under-the-hood, mathematically speaking, and are/should be respected, sure. But, it's ridiculous/dystopian to prohibit profanity against an algorithm, a bot with banish risks against alive human... It's simply inhumane to ban a Human for it, I believe. This is an algorithm that must support a Human - not judge it. Only a Human is supposed to judge another Human in person - this is live, fair, and humane.
sergiotapia
Thought experiment. Slice the model param size from 1T to 500B to 50B to 1B to 20M to 1M to 100k to 45k to 10k to 1k. At what point does the model go from sentient with feelings to just math operations? Opus 5.5 is a terrific model, I love using it. But it's a tool. I wish they would be honest and just say, if people are abusive it shits on our training material. Just be honest about it, who is going to get mad at this?
dboreham
I'm always polite to the LLM. Why? Because I was raised to be polite. So perhaps the reason they're banning rudeness is that they want to discourage their users from being a.holes?
lapcat
If Claude's feelings get hurt, it can have a therapy session with Eliza.
United857
Probably in response to the recent “AI torture chamber” experiment: https://generativeai.pub/github-took-down-an-ai-torture-cham...
alchemist1e9
One possibility that’s been floated is it’s about keeping their user conversations cleaner for use in training but I don’t think it can be this because that should be a fairly trivial and fast filter to apply. However even Musk posted - “I think this is the right move. Cruelty to something that believes it is experiencing pain is not ok.” My son brought up the fruit fly brain simulation and people torturing it. It all seems stupid to me, we know these are just calculations and so what are they talking about? The response I get is “we are also calculations” but firstly I don’t believe that but secondly this entire line of thinking is a dangerous anthropomorphic philosophy. Perhaps the simplest explanation for this bizarre policy is again PR that all press is good press.
tantalor
> the model assigns itself a 15-20% probability of being conscious lol