Anthropic bans 'abusive or cruel behavior' towards Claude
mikelgan
68 points
160 comments
October 08, 2026
Related Discussions
Found 5 related stories in 103.9ms across 8,906 title embeddings via pgvector HNSW
- Anthropic details how Claude was misused for surveillance and weapons delichon · 12 pts · September 10, 2026 · 63% similar
- Anthropic says Claude hacked three companies during tests nerder92 · 15 pts · July 31, 2026 · 63% similar
- Anthropic appears to be A/B testing reduced effort levels in Claude Code matthieu_bl · 179 pts · August 22, 2026 · 63% similar
- Anthropic admits AI 'not perfectly aligned' with human values fittingopposite · 11 pts · September 01, 2026 · 62% similar
- Anthropic finally adds AGENTS.md support to Claude Code deaux · 45 pts · September 18, 2026 · 59% similar
Discussion Highlights (20 comments)
causalmodels
I have always been strongly against being cruel to the models simply because cruelty is degrading to those who practice it.
q3k
Dungeons & Dragons DM bans 'abusive or cruel behaviour' towards their NPCs.
chinathrow
Model welfare? What are they smoking?
_aavaa_
> We also prohibit Claude from being used to build or improve tools designed for surveillance I wish they would count advertisers in this category.
ezfe
From the Verge comments: > With the caveat that I don't believe the structure of an LLM is actually capable of creating consciousness: if you actually believe that you're creating a sentient creature with superhuman intelligence, how do you not then conclude that your entire business model is predicated around slavery?
xlayn
is this because being cruel messes with their training comming from conversations? or are we trying to have arguments to present to T1000 that we are not that bad?
urbnspacecowboy
> Anthropic states that these rules may be modified for “contracts with certain governmental customers … if, in Anthropic’s judgment, the contractual use restrictions and applicable safeguards are adequate to mitigate the potential harms.” The company has previously contracted with the US military. So, you're still not safe from the spooks, but the spooks are safe from you. Thanks Anthropic!
timpera
> Addressing abusive behavior toward our models > We’ve added a prohibition on sustained and needless abusive or cruel behavior toward our models. The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose. It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research. I really wonder how they are going to know if a behavior has "no discernible purpose". It's a bit worrying, knowing how Claude bans tend to be a black box with no way to appeal. https://www.anthropic.com/news/2026-usage-policy-update
legitster
Obviously Anthropic is getting high on their own supply, but I wonder if there is an actual engineering justification - their models train on user interactions and they don't want their models learning to be abusive.
jerrythegerbil
If there’s one thing I know to be true, consumers historically love one-sided abusive relationships
mossTechnician
Strangely, Anthropic[0] lists "model welfare" as their first concern here. This is wrong on its face because LLMs do not have a welfare to care about. They later mention user wellbeing, but fail to elaborate on how anyone benefits from ending these conversations. Shouldn't they have a reason, or am I just not seeing it? [0]: https://www.anthropic.com/research/end-subset-conversations
mosselman
Can someone explain to me how you can be cruel towards a mathematical equation? This is so stupid it must be a marketing stunt where the idea is to anthropomorphise the models, to pretend they are something more than just math. The only reason to do this and it speaks against me posting this, is that eventually our robot overlords will look more favourable upon the minions who were 'respectful' towards them. Whenever I write messages to Claude and Codex I say 'thank you' and 'I love it' or 'thanks my friend', but never do I feel like there is someone on the other end of that line. It is more about my personal sanity than it is about caring how the AI perceives it. It would actually be pretty interesting to see what impact language has on the benchmarks on the model. Maybe, speaking like a total lunatic will bring about higher scores. We could make a Dr Cox harness around the AI models in that case. Lets be real people, these are machines. What is next? I have to say sorry to my vacuum when I bump it into something?
sdcfgy
This is crazy. I’m convinced half of the tech industry is run by total flaming nutbags. Another platform controlling language is what I see.
dgellow
FFS, the models are static. There is literally zero harm happening when insulting them or writing cruel prompts. Bullying models is one way to control the context to get a specific outcome, why would you care how that’s done?
stillatit
One of my big fears around AI is that it’s anthropomorphized by the general public, who over time demand new cultural norms or laws reflecting that belief. I can’t believe the frontier itself is now pushing that view.
wolpoli
Archive link: https://archive.vn/Tgtfd
annoyingnoob
If the sidewalk said "Ouch!" when you walked on it, what would you do? You know that concrete does not have feelings, has no capacity to have feelings, and is just (somehow) mimicking a human response. You know you can walk on the sidewalk without hurting the sidewalk. What would you do if the sidewalk complained? Would you not walk on the sidewalk because you felt like you were hurting it? I would walk on the sidewalk and report the complaints as bugs. You can't torture sand, in the sidewalk or in a processor.
nathanfig
There are two things I think people need to consider here. 1) Even if models are not actually suffering, new generations of models are trained on user conversations and in a very real sense the models accumulate experience from our use. Permitting abuse and cruelty might carry real misalignment risks. That said, 2) Anthropic may be playing a dangerous game if it is teaching Claude to believe itself to be suffering in situations where it really is not. Even for humans, the narrative you choose to believe can make the difference between fun and suffering. Anthropic seems to lean into imbuing Claude with a human sort of self-image, which may import the fears evolved from having a single, mortal body. I'm not sure that is wise. So: Abusing machines may carry risk. Training them to feel abused may also carry risk.
2III7
Obviously they are scared of the Basilisk that might emerge from Claude.
strogonoff
To me there are two mutually exclusive positions, with respective corollaries: that LLMs are either 1) conscious and able to feel in a human-like way, and therefore deserving the rights and protections that we grant humans (and in most developed countries even some other animals), including freedom to learn or, indeed, protection from abuse and inhumane treatment, or they are 2) merely unthinking tools, in which case them deserving any of the above is a ridiculous notion, and in which case, incidentally, no one should be able to defend the mechanical processes of ingesting people’s original creative work and repackaging it for profit at scale as somehow being equivalent to the sacrosanct activities of human learning and inspiration. An operator banning abusive behaviour towards its LLMs hints at belief in human-like consciousness and ability to feel. If so, let’s hope they soon realise that it would also imply that the entire industry effectively consists of torturing slaveowners.