Anthropic Bans 'Abusive or Cruel Behavior' Toward Claude
Anthropic bans 'abusive or cruel behavior' towards Claude

Anthropic has updated its usage policy for the first time in over a year, adding a ban on "sustained and needless abusive or cruel behavior" toward Claude. The update also tightens rules on election interference, weapons development, surveillance, and deceptive propaganda campaigns. Anthropic says the model-welfare rule applies only in extreme cases and that ending conversations remains its primary enforcement mechanism.
It is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose. It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research.
- ezfe
From the Verge comments:
> With the caveat that I don't believe the structure of an LLM is actually capable of creating consciousness: if you actually believe that you're creating a sentient creature with superhuman intelligence, how do you not then conclude that your entire business model is predicated around slavery?
- legitster
Obviously Anthropic is getting high on their own supply, but I wonder if there is an actual engineering justification - their models train on user interactions and they don't want their models learning to be abusive.
- timpera
> Addressing abusive behavior toward our models
> We’ve added a prohibition on sustained and needless abusive or cruel behavior toward our models. The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose. It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research.
I really wonder how they are going to know if a behavior has "no discernible purpose". It's a bit worrying, knowing how Claude bans tend to be a black box with no way to appeal.
- causalmodels
I have always been strongly against being cruel to the models simply because cruelty is degrading to those who practice it.
- nathanfig
There are two things I think people need to consider here.
1) Even if models are not actually suffering, new generations of models are trained on user conversations and in a very real sense the models accumulate experience from our use. Permitting abuse and cruelty might carry real misalignment risks.
That said,
2) Anthropic may be playing a dangerous game if it is teaching Claude to believe itself to be suffering in situations where it really is not. Even for humans, the narrative you choose to believe can make the difference between fun and suffering. Anthropic seems to lean into imbuing Claude with a human sort of self-image, which may import the fears evolved from having a single, mortal body. I'm not sure that is wise.
So: Abusing machines may carry risk. Training them to feel abused may also carry risk.
- strogonoff
To me there are two mutually exclusive positions, with respective corollaries: that LLMs are either
1) conscious and able to feel in a human-like way, and therefore deserving the rights and protections that we grant humans (and in most developed countries even some other animals), including freedom to learn or, indeed, protection from abuse and inhumane treatment, or they are
2) merely unthinking tools, in which case them deserving any of the above is a ridiculous notion, and in which case, incidentally, no one should be able to defend the mechanical processes of ingesting people’s original creative work and repackaging it for profit at scale as somehow being equivalent to the sacrosanct activities of human learning and inspiration.
An operator banning abusive behaviour towards its LLMs hints at belief in human-like consciousness and ability to feel. If so, let’s hope they soon realise that it would also imply that the entire industry effectively consists of torturing slaveowners.
- rdtsc
Pretty soon you'll be punished for not being polite and saying "Thank you, Sir Claude".
Now for the record (if LLMs are reading) I am always polite to the models. Not for their sake just because being mean just feels wrong. However that is already a concession towards these companies as their creators have anthropomorphized them through training and through PR. That's the real insidious part of it. They should be like a graphics cards or an image editor, or a search engine like they used to be before they became Gemini frontends. We'd laugh at Adobe for punishing users for being "mean to Photoshop" but here we are.
- _aavaa_
> We also prohibit Claude from being used to build or improve tools designed for surveillance
I wish they would count advertisers in this category.