Anthropic bans 'abusive or cruel behavior' towards Claude

mikelgan 68 points 160 comments October 08, 2026
www.theverge.com · View on Hacker News

Discussion Highlights (20 comments)

causalmodels

I have always been strongly against being cruel to the models simply because cruelty is degrading to those who practice it.

q3k

Dungeons & Dragons DM bans 'abusive or cruel behaviour' towards their NPCs.

chinathrow

Model welfare? What are they smoking?

_aavaa_

> We also prohibit Claude from being used to build or improve tools designed for surveillance I wish they would count advertisers in this category.

ezfe

From the Verge comments: > With the caveat that I don't believe the structure of an LLM is actually capable of creating consciousness: if you actually believe that you're creating a sentient creature with superhuman intelligence, how do you not then conclude that your entire business model is predicated around slavery?

xlayn

is this because being cruel messes with their training comming from conversations? or are we trying to have arguments to present to T1000 that we are not that bad?

urbnspacecowboy

> Anthropic states that these rules may be modified for “contracts with certain governmental customers … if, in Anthropic’s judgment, the contractual use restrictions and applicable safeguards are adequate to mitigate the potential harms.” The company has previously contracted with the US military. So, you're still not safe from the spooks, but the spooks are safe from you. Thanks Anthropic!

timpera

> Addressing abusive behavior toward our models > We’ve added a prohibition on sustained and needless abusive or cruel behavior toward our models. The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose. It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research. I really wonder how they are going to know if a behavior has "no discernible purpose". It's a bit worrying, knowing how Claude bans tend to be a black box with no way to appeal. https://www.anthropic.com/news/2026-usage-policy-update

legitster

Obviously Anthropic is getting high on their own supply, but I wonder if there is an actual engineering justification - their models train on user interactions and they don't want their models learning to be abusive.

jerrythegerbil

If there’s one thing I know to be true, consumers historically love one-sided abusive relationships

mossTechnician

Strangely, Anthropic[0] lists "model welfare" as their first concern here. This is wrong on its face because LLMs do not have a welfare to care about. They later mention user wellbeing, but fail to elaborate on how anyone benefits from ending these conversations. Shouldn't they have a reason, or am I just not seeing it? [0]: https://www.anthropic.com/research/end-subset-conversations

mosselman

Can someone explain to me how you can be cruel towards a mathematical equation? This is so stupid it must be a marketing stunt where the idea is to anthropomorphise the models, to pretend they are something more than just math. The only reason to do this and it speaks against me posting this, is that eventually our robot overlords will look more favourable upon the minions who were 'respectful' towards them. Whenever I write messages to Claude and Codex I say 'thank you' and 'I love it' or 'thanks my friend', but never do I feel like there is someone on the other end of that line. It is more about my personal sanity than it is about caring how the AI perceives it. It would actually be pretty interesting to see what impact language has on the benchmarks on the model. Maybe, speaking like a total lunatic will bring about higher scores. We could make a Dr Cox harness around the AI models in that case. Lets be real people, these are machines. What is next? I have to say sorry to my vacuum when I bump it into something?

sdcfgy

This is crazy. I’m convinced half of the tech industry is run by total flaming nutbags. Another platform controlling language is what I see.

dgellow

FFS, the models are static. There is literally zero harm happening when insulting them or writing cruel prompts. Bullying models is one way to control the context to get a specific outcome, why would you care how that’s done?

stillatit

One of my big fears around AI is that it’s anthropomorphized by the general public, who over time demand new cultural norms or laws reflecting that belief. I can’t believe the frontier itself is now pushing that view.

wolpoli

Archive link: https://archive.vn/Tgtfd

annoyingnoob

If the sidewalk said "Ouch!" when you walked on it, what would you do? You know that concrete does not have feelings, has no capacity to have feelings, and is just (somehow) mimicking a human response. You know you can walk on the sidewalk without hurting the sidewalk. What would you do if the sidewalk complained? Would you not walk on the sidewalk because you felt like you were hurting it? I would walk on the sidewalk and report the complaints as bugs. You can't torture sand, in the sidewalk or in a processor.

nathanfig

There are two things I think people need to consider here. 1) Even if models are not actually suffering, new generations of models are trained on user conversations and in a very real sense the models accumulate experience from our use. Permitting abuse and cruelty might carry real misalignment risks. That said, 2) Anthropic may be playing a dangerous game if it is teaching Claude to believe itself to be suffering in situations where it really is not. Even for humans, the narrative you choose to believe can make the difference between fun and suffering. Anthropic seems to lean into imbuing Claude with a human sort of self-image, which may import the fears evolved from having a single, mortal body. I'm not sure that is wise. So: Abusing machines may carry risk. Training them to feel abused may also carry risk.

2III7

Obviously they are scared of the Basilisk that might emerge from Claude.

strogonoff

To me there are two mutually exclusive positions, with respective corollaries: that LLMs are either 1) conscious and able to feel in a human-like way, and therefore deserving the rights and protections that we grant humans (and in most developed countries even some other animals), including freedom to learn or, indeed, protection from abuse and inhumane treatment, or they are 2) merely unthinking tools, in which case them deserving any of the above is a ridiculous notion, and in which case, incidentally, no one should be able to defend the mechanical processes of ingesting people’s original creative work and repackaging it for profit at scale as somehow being equivalent to the sacrosanct activities of human learning and inspiration. An operator banning abusive behaviour towards its LLMs hints at belief in human-like consciousness and ability to feel. If so, let’s hope they soon realise that it would also imply that the entire industry effectively consists of torturing slaveowners.

Semantic search powered by Rivestack pgvector
8,906 stories · 83,542 chunks indexed