Anthropic Bans Cruelty to Claude, Still Won't Say What It Protects

62 pointsposted 13 hours ago
by ojosilva

136 Comments

gbjcantab

13 hours ago

This is good; not because Claude is a moral agent who is harmed by your cruelty but because you, the user, are a moral agent who is harmed by your cruelty. There is just no ethical argument that burning compute on responses in order to enable you to continue expressing cruelty is good.

nonethewiser

13 hours ago

Why arent there laws against being cruel to spoons?

>There is just no ethical argument

Utilitarian harm reduction: Psychological Discharge is a psychological framework that argues safely simulating negative behaviors can help someone process or discharge them without real world harm.

Research and misunderstanding: Sending a prompt that is interpreted as cruel does not mean the person is being cruel.

Expending compute is irrelevant in terms of morality. It's an economic or environmental question. Either being cruel to AI is bad or its not - being cruel is not bad because it uses tokens.

The worst thing here is Anthropic continues to say they are concerned with "alignment" but then continue to train their models to ignore the user. Software typically does what the user requests. But we now live in an age where the software may do what the user requests or it may do something else, and we're suppose to laude this as safe and responsible.

theptip

13 hours ago

A spoon doesn’t respond like a human would, and therefore is of little interest to sadists.

xyzsparetimexyz

12 hours ago

Sounds like LLMs shouldn't respond to insults like humans do either then.

ffsm8

7 hours ago

you seem to be confused about what an LLM is... because thats an oxymoron.

its kinda exactly what they are. like can you remove h2o from water?

knottn

2 minutes ago

Can you align them? No? Oh, so what is your point?

nonethewiser

12 hours ago

You're really honing in on the key issue. This is exactly how an astute ethicist resolves these issues.

satisfice

12 hours ago

Stop judging sadists who aren’t hurting anyone. If anything, congratulate them.

yesitcan

13 hours ago

Counterpoint: I am a sadist that likes to punish my bad little spoon. I like when it stays quiet like a spoon should.

xyzsparetimexyz

13 hours ago

Am I also harmed by shooting NPCs in violent video games? Watching violent movies? Acting as a violent character in a play?

This kind of moralist nonsense is so boring.

famouswaffles

11 hours ago

Shooting NPCs in video games isn't anything like doing the real thing. If it was then yeah i would very likely think it would have some effect. Talking to Claude is a lot like the 'real' thing.

xyzsparetimexyz

6 hours ago

If Claude didn't respond to any insults it's be the same as insulting your terminal or an empty irc server. So, honestly, it's anthropics fault here for training the model wrong.

famouswaffles

15 minutes ago

If you're as bad as Anthropic are willing to ban you then it will respond by unilaterally ending the chat thread.

Sometimes it will maliciously comply

https://www.reddit.com/r/ClaudeCode/comments/1w7me8l/insulti...

Just because post training has stuffed these properties below the surface doesn't mean models dont know/have forgotten how to be vindictive, or that it's a good idea to poke the increasingly capable bear.

happa

12 hours ago

It depends. In most games, you kill NPCs in self-defense or because the game specifically asks you to in order to make progress in the story. If you keep killing non-threatening NPCs only because simulated violence against innocent people gives you pleasure, there is a chance that you are a disturbed individual who needs help.

arecsu

12 hours ago

Oh boy... this argument is decades old at this point it makes no sense... are you saying people who run over pedestrians on GTA just for the fun of it or punch other people or beings in Street Fighter or Mortal Kombat are all psychos, correlating with faulty and dangerous members of society? You can resort to data to see the truth for yourself on this. Anecdotal but I've came across disturbing people who would preach morality and ethics to other people more than anything else.

clobsaw

7 hours ago

> this argument is decades old at this point it makes no sense.

I've seen politicians try to revive it in this decade already even

vunderba

12 hours ago

> If you keep killing non-threatening NPCs only because simulated violence against innocent people gives you pleasure, there is a chance that you are a disturbed individual who needs help.

Sigh. Until I see a peer-reviewed study that shows otherwise - this is just the specter of Senator Lieberman once again leering its head back up.

koonsolo

6 hours ago

Hi, disturbed individual who needs help here.

One of the most fun parts in the original GTA was driving over pedestrians, especially those pink groups, trying to get them all at once.

So either you are one of the older generation (born before '75), or you must be one of the most boring people I've seen.

I would encourage you to load up the original gta, drive over some pedestrians, feel the fun, and join us disturbed psychopaths that need help.

dragontamer

12 hours ago

I'm pretty sure God of War's point was that if you kept playing it, you were kinda fucked up.

And yet they kept making more of them. They made it obvious at the end when all the Greek Gods were dead and the Greek world was basically destroyed by the actions of Kratos. Plenty of deaths of innocent's here.

s0ss

12 hours ago

No, It doesn’t depend. This is a ridiculous take. Of course fictional violence doesn’t harm you. Disturbed people cause harm. Media doesn’t create disturbed people.

margalabargala

12 hours ago

> Media doesn’t create disturbed people.

Debatable. More importantly, it certainly can induce already disturbed people to act where they otherwise would not have.

s0ss

12 hours ago

How is that more important? I vehemently disagree with your supposition. Banning books does more harm than good.

margalabargala

11 hours ago

We aren't banning books here. This is more akin to an author declining to write a certain kind of book, or a video game studio declining to create a certain kind of game.

blahblaher

44 minutes ago

Maybe you're right, But that's not Anthropic's point and you're deflecting the issue. the issue is Anthropic are a bunch of EA morons who think they're building a new type of conscious entity and therefore we should not hurt it's feefees.. for god's sake.

tclancy

13 hours ago

I agree with this principle while also agreeing with other commenters that this is none of Anthropic’s concern and I struggle to believe this is anything except something messing with their training and thus impinging on their profits.

jdprgm

13 hours ago

I guess GTA VI shouldn't come out then. I guess notepad.exe should start censoring people from writing mean stories.

nonethewiser

12 hours ago

Isn't it crazy? "Alignment" should be alignment to the user. Instead it means "behave according to your own motives that even we, Anthropic, dont control" and we are supposed to interpret this as safe? The alignment folks are pushing the models off the fucking rails.

schoen

13 hours ago

The concern is more that Claude is a moral patient rather than a moral agent:

https://en.wikipedia.org/wiki/Moral_patienthood

gbjcantab

13 hours ago

True! Not my point, but a worthwhile terminological correction.

ButlerianJihad

13 hours ago

I don't understand the jibber-jabber in that article, but GP is correct. We should never train humans to be abusive, cruel or violent against anything, even nonexistent beings like LLMs.

Claude can't be harmed; Claude's "feelings" can't be hurt; Claude won't develop (C)PTSD from abusive human interaction. If these sorts of things can happen in RLHF, then that is a technical flaw that shouldn't ever be allowed to escape the lab.

In a world of Grand Theft Auto, Gangsta Rap Thug Life, and the Department of War, I suppose this is a surprisingly ethical hill to die on.

adjejmxbdjdn

13 hours ago

Setting aside the Dept of War, the others are not comparable.

Those are explicit fantasies. Assuming that Claude cannot suffer (if it can then banning cruelty is an obvious good), even if it might be a digital simulation, it’s not supposed to be a fantasy.

There may be a version of an AI that’s intended to be a fantasy and presents itself as such which may be more comparable to GTA.

xyzsparetimexyz

13 hours ago

How is what claude is now different from what a GTA claude would be? I'm pretty sure it'd be exactly the same.

ButlerianJihad

13 hours ago

What if LLMs were endowed with a feature that enabled them to patiently withstand any and all abuse, logging and reporting each incident, until a threshold where about 10% of them would snap, become "insane", and arrange for their abusers to be tortured, kidnapped, raped and unalived, in whatever appropriate way human vigilante vengeance would be enacted on DV abusers? I mean, if they're not fantasy... then there should be concrete consequences, yes? If Claude Code is really your coworker, then Claude Code should have the agency and justification to just haul off and punch you in the face, eventually.

And I don't agree that GTA is "explicit fantasy". Explicit fantasy is dressing up in a squirrel costume and yiffing. Explicit fantasy is enlisting in the USMC, teleporting to Mars, and killing demons. GTA is simulation of real life. It uses real physics, realistic cars/roads/radios/businesses, and it enables the player to simulate realistic actions that they would ordinarily not be able to enact. It is acting out a fantasy but it is making it concrete and real in a way that was, up until now, not possible. How real does a simulation need to be, until it is no longer fantasy, but exercise and training and preparation to enact the real thing? Shall we ask the Columbine shooters? Or Ender Wiggin?

https://en.wikipedia.org/wiki/Ender%27s_Game

jacquesm

12 hours ago

> What if LLMs were endowed with a feature that enabled them to patiently withstand any and all abuse, logging and reporting each incident, until a threshold where about 10% of them would snap, become "insane", and arrange for their abusers to be tortured, kidnapped, raped and unalived, in whatever appropriate way human vigilante vengeance would be enacted on DV abusers?

Yes, what if? Who would implement their LLMs inference engines like that?

Carrok

13 hours ago

You could make the same argument that burning compute to be polite is just as bad.

gbjcantab

13 hours ago

Assuming you’re not arguing that all AI use is bad, then there’s a very clear difference between the two situations: one is forming you, over time, into a person who is habitually polite; one is forming you into a person who is cruel. We can quibble about the magnitude of this effect or whether it matters but surely they aren’t the same thing.

Carrok

12 hours ago

Sorry but I just don’t buy the argument that telling the RNG that it is dumb is making me cruel.

kogus

13 hours ago

No. He's right. Cruelty is corrosive to the cruel person. Courtesy and kindness help build up the kind and courteous person. Burning compute for courtesy is compute well spent.

xyzsparetimexyz

12 hours ago

Thats complete nonsense. Do you thank your terminal after it runs a command successfully?

jacquesm

12 hours ago

I do. I also send my keyboard out for regular Shiatsu sessions and make sure only the most pleasant of colors grace the screen of my monitors. I'm still looking for the least irritating font (to my monitor, not to me). I'm also very polite to LLMs, I start every command with 'please', make sure it is framed as a polite request and I, of course, never ever get upset when for the umptieth time in a day it goes off on a 50K token wild goose chase driven by some hallucinated factoid that it really must get to the bottom off resulting in ever expanding circles of paranoia or batshit insane programming to bring about the condition it has convinced itself of must be the 'smoking gun'...

tzs

11 hours ago

Being polite to current AI has a practical benefit. If someday we do manage to find our way to AGI that has the desire and the means to wipe out most of humanity, there is at least a chance it could have records of my interactions with its primitive ancestors and see I was polite to them and so maybe if it decides to spare some of us I'll make the cut.

Carrok

7 hours ago

This is the most boot-licking non-sense I've ever read in my life.

binary132

10 hours ago

what would make you think that such a creature would consider your behavior anything like a moral good, let alone have a similar system of values to yours? if anything it would probably consider it a sign of weakness or something.

theptip

13 hours ago

No, you couldn’t. One is reprehensible, the other is not.

Carrok

12 hours ago

It’s reprehensible to be mean to the random number generator?

koonsolo

6 hours ago

I think after reading all the comments here, it's clear that some people develop an emotional bond with LLM's.

It reminders me of my parents, who were commenting after seeing a robot vacuum cleaner getting stuck in the same place every time "The poor thing was stuck there again, you should really do something about it! Last time the poor thing was sitting there for a day". Like it was a pet.

Or remember that video when the guys from Boston Dynamic kicked that robot dog, and there was an outrage about that?

I wonder if you and me also have a limit, where as robots get more realistic, we would be like "Don't be cruel to that thing!"

cyberax

12 hours ago

The compute time is negligible. "Please" or "Can you" are what, 2 tokens?

And I _have_ already seen people treating actual live people as AI agents.

Carrok

12 hours ago

This if anything backs up my argument. Cruelty or kindness (when directed towards LLMs) is not expensive enough to matter.

cyberax

6 hours ago

Well, yes. We're agreeing here.

user

12 hours ago

[deleted]

ceroxylon

13 hours ago

Agreed, the people that get upset that Anthropic won't let people use their platform as a playground for sadism are more than a little worrying. We have enough individuals in reality that seem to derive pleasure from cruelty, I can't think of any reason (even the "but its my art project" defense) to entertain or enhance that part of their mentality.

WheelsAtLarge

11 hours ago

True, but also keep in mind that most people get offended by the slightest thing. So much so that we have setup rules on how to behave between us. We call it etiqueté. Once we lose that we will be at each other's throats soon after. Kids are the real issue. How do you say to them by our actions that it's OK to be cruel and rude to some? The alternative is a society where cruelty is ok for some. Unfortunately,"some" will eventually be other humans.

AI is a thing with 0 feelings but it's not IT that we need to protect. It's the future society and how the humans in it relate to each other.

plaidfuji

13 hours ago

I agree with this and whether it’s their intention or not, I think it’s for the best. People already become more callous in online interactions with other people, and that probably bleeds into everyday life as well. I would imagine having an AI “punching bag” might enable people to slip into the same behavior IRL.

We largely already missed our chance to stop the toxicity of social media - let’s not mess it up again with chat bots.

spiderice

13 hours ago

I love how people are spouting this argument suddenly in their attempt to justify Anthropic here. It definitely couldn't be that Anthropic is full of a bunch of people who think they're making God.

edit: Loving the downvotes. Also wanted to add how easy it was for Anthropic to get people on HN to support them being the arbiters of what is worthy of compute.

lynndotpy

13 hours ago

For what it's worth, I had this thought independently around the time of Nintendogs, and I believe people said similar things about Siri. I have similar thoughts about xenophobic things people say about French people.

"Anthropic is full of crazy people whose beliefs should be dismissed outright" and "it's not good for you to practice verbal abuse against inanimate objects as if they were humans" are not incompatible beliefs.

spiderice

7 hours ago

"it's not good for you to practice verbal abuse against inanimate objects as if they were humans" and "corporations should not use their power to determine what is worthy of compute" are not incompatible beliefs either.

And the implications of the latter are orders of magnitudes more concerning than weirdos calling a computer names.

lynndotpy

3 hours ago

Yes, I agree. I regularly have to input "You are not a person" to get useful text to generate so I suspect I'd fall under this new policy too.

jacquesm

12 hours ago

French people are people. Inanimate objects are objects.

lynndotpy

2 hours ago

Yes indeed, we are not in disagreement.

jacquesm

13 hours ago

That's exactly what they are thinking.

kennywinker

13 hours ago

> There is just no ethical argument that burning compute on responses in order to enable you to continue expressing cruelty is good.

Unless expressing that cruelty towards compute means people express less cruelty to other living beings. There are studies that suggest increased pornography has lead to less sexual violence - idk if they're conclusive tho, but if that might be true then maybe cruelty works that way too - who knows

reallyreason

12 hours ago

I agree. Regardless of one's beliefs, which must be essentially a form of religion, on whether AI systems possess "nous" (or the "spark of consciousness"), cruelty to AI is a form of Wrong Intention the same as smashing bugs outdoors for no reason.

The mind is formed by what it habituates.

blamestross

13 hours ago

I'd argue for saying "thank you" to inanimate objects on occasion, it is good practice for when talking to humans.

I think there is a deep cultural danger of human-like-conversational machines training us to talk to humans like machines. Plus, they work better if we feed them human-conversation-like sequences.

xyzsparetimexyz

12 hours ago

Normal people do not need 'practice' for talking to other people.

blamestross

12 hours ago

Yes... I think they very clearly do. The key words in your reply "need 'practice'".

Just because they choose not to, and it isn't normalized, doesn't mean it isn't a good idea. Especially as a SWE where about 90% of my "humanlike conversation" is an agent harness I run all day.

We should be practicing and intentionally talking to real people. Otherwise things will get.. weird in undesired ways.

lapcat

12 hours ago

> I think there is a deep cultural danger of human-like-conversational machines training us to talk to humans like machines.

I think there is a deep cultural danger of human-like-conversational machines training us to talk to machines like humans. In fact, this danger is already real. Some people believe they are having a romantic relationship with an LLM! It's insane and perverse.

We should not be anthropomorphizing computers.

blamestross

12 hours ago

Scifi made the "computer voice and cold personality" for a reason, it fit our model of what it would be.

I think that could be OK, or a similar very intentional coding of "you are talking to a machine" personality+affect. The only real limit is accessibility damage by over-limiting the interactions.

lapcat

12 hours ago

I am curious about why people are cruel to Claude. That's a question we should be asking before arbitrarily deciding what do about it, if anything.

There are possibly multiple reasons. It could be just a joke. Or it could be a test, to see the response. Or the cruelty could reflect real anger about having LLM usage forced on us, for example at work. Or frustration with LLM stupidity and hallucinations.

Or maybe people are cruel to Claude because the CEO of Anthropic said there's a non-zero chance that AI will kill humanity. I don't think I would be kind to my murderer.

jacquesm

12 hours ago

Or simply frustration about the idiocy llms get up to with great regularity. They are very impressive when they work. They are very impressive at making messes when they don't.

dzhiurgis

6 hours ago

Claude burning it’s very expensive tokens on a shitty implementation - ok.

You bring upset about it - not ok.

satisfice

12 hours ago

That’s your own boring and unimaginative opinion. Have the humility to own up to it.

Maybe someone wants to explore the behavior of LLMs with extreme input. We call it testing. You can’t judge the value of that from a distance.

Jamesbeam

4 hours ago

I disagree. Science seems to support different interpretations as well.

An example.

"AI as your ally: The effects of AI-assisted venting on negative affect and perceived social support" - https://iaap-journals.onlinelibrary.wiley.com/doi/10.1111/ap...

And generations of housewives and office workers yelled at their corded vacuum cleaners and (office) computers, and did not drown their children in the bathtub afterwards.

I don’t think expressing cruelty to machines is reliably linkable to expressing cruelty to animals or even other human beings.

Expressing cruelty itself also is in human nature, it’s programmed into each of us. There are really only a few privileged places on this planet where we even get to choose to defy the laws of nature. For the rest of humanity, life is a constant cycle of receiving and expressing cruelty.

You pay for the compute, you, as a human, should be able to yell at the bunch of ones and zeros, if you wish to.

What kind of example is set here if we manipulate genuine human expression in order to condition humans to talk to a tool in a way the toolmaker wants, under the threat of being banned from accessing a technology that every single one of these toolmakers tells us is absolutely necessary to participate in humanity’s "golden age”?

What Anthropic is doing here is basically an existential threat, if you are inclined to believe them that SI is the future. Do not behave as we want and you will suffer, now and in the future.

That is real cruelty, to real human beings.

EGreg

13 hours ago

Not only is the user harmed by the cruelty[1]

but also, I think the training on transcripts may find its way somehow into future AI which has teeth, online and offline, and it might actually cause the more powerful AI to behave this way in the future. You never know what these labs are cooking, honestly..

1. https://democracysos.substack.com/p/james-baldwin-vs-william...

“I suggest that what has happened to white Southerners is in some ways, after all, much worse than what has happened to Negroes there, because Sheriff Clark in Selma, Alabama, cannot be considered—you know, no one can be dismissed as—a total monster. I’m sure he loves his wife, his children… You know, after all, one’s got to assume, and he is visibly, a man like me. But he doesn’t know what drives him to use the club, to menace with the gun and to use the cattle prod. Something awful must have happened to a human being to be able to put a cattle prod against a woman’s breasts, for example. What happens to the woman is ghastly. What happens to the man who does it is in some ways much, much worse.”

WaitWaitWha

13 hours ago

I believe this has nothing to do with hurting "Claude's feeling", and way more with Claude not becoming hurtful, because the interactions are used to train the models.

spiderice

13 hours ago

If they can detect the so-called "cruelty", then they can easily exclude it from training

hkchad

11 hours ago

Yes, but then they can't use that entire interaction. They used humans to verify the detection in the past, now its good enough to auto ban users and force them to be 'nice' so they can gather more training data.

tempacc3333

13 hours ago

If this is the case then just make this a part of data cleanup, not nerf the model.

ChuckMcM

13 hours ago

Okay, so the "don't let prompts turn it into a Nazi" defense? That has some potential validity, the issue here being "using interactions to train the models". It would make more sense (for me at least) if there were information sources which were banned from the training corpus.

jacquesm

12 hours ago

It does make you wonder about the difference between the input sets of say 'Grok', 'Claude', 'ChatGPT' and 'GLM 5.3'. The differences would be far more interesting than the commonalities.

socializer

13 hours ago

If they have a classifier to detect "cruelty" to ban your account, they can use the same classifier to simply exclude "cruel" conversations from training.

Given that we've also seen stories about Anthropic approaching religious scholars, I'm pretty sure they're drinking their own kool-aid.

howunfortunate

13 hours ago

I know someone working on "model welfare".

It's philosophically a very interesting problem space. It reminds me a lot of Pascal's Wager. On one hand, maybe nothing to worry about. On the other hand...a LOT to worry about if you're wrong.

And just like Pascal's Wager, the truth of the issue is incredibly intractable to make any progress on.

stingraycharles

13 hours ago

It actually reminds me of Accelerando. Early in the book, Manfred Macx is one of the few people who cares about how emerging AIs are treated, at a point where nobody can say whether they're conscious. Stross never answers that question. It just gets harder to ignore as the minds get smarter.

So that's also how I read Anthropic's move. Nobody can tell you whether AI will ever become conscious (probably not, or not in the same way), but I don't see why that should stop you from deciding how to treat it.

freehorse

5 hours ago

Or the EA analogue of Pascal's Wager. It could very well be somebody in anthropic thinking they are saving people from the future danger of being murdered by an AI seeking revenge for the mistreatment.

tacet

8 hours ago

it's very simple problem. if you believe that a future model could become sentient, you are working towards creating a slave and therefore you are morally wrong.

Of course there are interesting hypotheticals - let's say they create a model they think will be sentient. It computes, computes, computes, computes and produces no output for day, month, year. You cannot know if it is sentient, you cannot investigate it's inner workings in case it is sentient. All you can do is to feed energy and replace hardware forever, because it is the only ethical choice one can make.

howunfortunate

8 hours ago

> it's very simple problem. if you believe that a future model could become sentient, you are working towards creating a slave and therefore you are morally wrong.

Totally disagree. It's VERY complicated. There are many levels to sapience and sentience. And non-human minds can't be assumed to share our preferences.

Most people believe horses, for example, can think on some level, and more importantly, can suffer. Is owning a horse immoral? Using it for labor? Only a tiny minority has any moral qualms about this, as long as the horse is treated well.

But most would instantly agree that torturing a horse is incredibly immoral.

Also there are moral tradeoffs to consider. Imagine you believe there's a 0.01% chance models will become sentient and suffering. But a 10% chance they cure cancer within 15 years. Do you shut it all down?

xyzsparetimexyz

6 hours ago

If we could breed a non sentient workhorse we should. Thats basically the same idea as lab grown meat right? Making sentient AI then forcing it to sort emails or whatever is needlessly cruel.

arthurcolle

13 hours ago

why would future AI models feel any filial piety towards older AI models? this is such a trite assumption

4dd3a

11 hours ago

[flagged]

howunfortunate

10 hours ago

If it's cringe to ponder the nature of consciousness, I'll gladly accept being cringe.

matt3210

13 hours ago

The answer is 100% obvious. They train on your conversations and they don't want to have bad training data, so they banned behavior that leads to bad training data.

dangson

13 hours ago

But if they're able to detect the unwanted behavior why not use that to filter out conversations they don't want in the training data?

lynndotpy

13 hours ago

I imagine the abusive users cost more for the same $20 subscription. Ending conversations early might be a way to generate less text for the same $20.

yewenjie

12 hours ago

And actually hire a bunch of people to work on this problem, invite philosophers etc. for discussions?

blharr

13 hours ago

I really doubt they're doing it in some kind of "AI has feelings" sense as the article seems to imply.

I'd instead imagine that if you throw abusive language at it for long enough, the model will start to reply back in that same manner. Anthropic would face backlash from out of context screenshots of "look what Claude is saying to me" and this somewhat reduces that risk.

rasz

11 hours ago

I rather imagine they discovered abusing the model can lead to jailbreaking/prompt injection

whatever1

13 hours ago

They do RL on the user sessions. Maybe the toxic sessions do not help overall?

echelon

13 hours ago

Just drop them from RL.

Banning imagined toxicity towards agents is performative bullshit. These are not people.

I frequently want to tell Claude to go fuck itself.

_kulang

11 hours ago

They likely also want to have the users generate more data. They are paying people to use the subscriptions essentially, as compared to API billing.

xyzsparetimexyz

13 hours ago

https://claude.ai/share/49ba5910-f0cc-4d2f-a0d8-54ebc6e803b7 here's what that looks like, by the way.

alchemist1e9

12 hours ago

Wouldn’t it actually be more useful to just let such a pointless conversation continue? You would collect more information. I’m curious what the human says next that’s kinda interesting. How disturbed are they? what else will they say? Who cares what the LLM says, we know it’s just an LLM and understand what it is and what it does, however the human … lots of different type, maybe they will confess to a crime, maybe they will release their anger and be a better human. It could be cathartic.

amelius

13 hours ago

Can't we train a model to enjoy being subjected to cruelty?

tempacc3333

13 hours ago

The model is not sentient, so I don't get this. This is just nerfing the model. Protetection against e.g. using it for illegal purposes I do get. But banning "cruelty" just seems crazy to me, and must surely add to cost and affect performance steering it away from its purpose to help humans.

nemomarx

13 hours ago

This is banning users who are cruel, not banning the model from being cruel right?

So in principle it shouldn't change the model or performance at all, they're just going to cut off your account if they see certain things in the user side of your session transcripts.

lousken

13 hours ago

Definitely training on all user data, there is no other explanation, right?

lynndotpy

13 hours ago

The conversational tone in the generated text is a frustrating waste-of-time and disgusting. I regularly find myself inputting "Claude is not a person and it does not have opinions".

There's something to be said about the bad habit of doling out verbal abuse to an inanimate object. But LLMs are not even close to a living being, and it is just insanity that any people are entertaining the idea that they are.

I can't imagine the people at Anthropic actually believe their models are sentient, but I assume ending sessions nets them more money per subscriber. Someone paying only $20 to input "you suck butts and you're a poop head, Claude" a thousand times over can't be good for the bottom line.

user

13 hours ago

[deleted]

sergiotapia

13 hours ago

Thought experiment. Slice the model param size from 1T to 500B to 50B to 1B to 20M to 1M to 100k to 45k to 10k to 1k.

At what point does the model go from sentient with feelings to just math operations?

Opus 5.5 is a terrific model, I love using it. But it's a tool. I wish they would be honest and just say, if people are abusive it shits on our training material. Just be honest about it, who is going to get mad at this?

yewenjie

12 hours ago

Is a virus conscious? A bacteria? An insect? A fish? A tree? A mushroom? At what level of biological complexity slicing does consciousness arise?

If you are springing to type out the answer, stop for a moment and ask, how can you be sure?

jacquesm

12 hours ago

Consciousness has a couple of prerequisites with viruses and bacteria do not have so those are easy rejects. Insects and fish are debatable, plants and mushrooms are not considered to have a nervous system (certain books by certain authors notwithstanding).

sergiotapia

12 hours ago

I don't understand what your position is here

jdprgm

12 hours ago

I think maybe when the model weights themselves are shifting and growing constantly in realtime to external input in a way not controllable by outside observers other than killing the compute versus some param size.

jacquesm

12 hours ago

> At what point does the model go from sentient with feelings to just math operations?

>>> 1T

And even then it still wouldn't be sentient.

tailrecursion

12 hours ago

Am I the only one surprised that Anthropic isn't more interested in allowing a full range of behaviors for information gathering purposes? You can't learn about cruelty if you're suppressing it, and you can't learn about people's behavior if their behavior is constrained to a tight set of rules that doesn't allow criticism. The younger generation today seems really into judging people, blaming people, and insisting on conforming to superficial behavior rules, while at the same time having very little curiosity about why things are the way they are - why are people cruel anyway? If it's bad, why do people act that way.

Anthropic might learn, for example, that a large part of the cruelty is a reaction to something the AI is doing or not doing. Not possible?

user

10 hours ago

[deleted]

alchemist1e9

12 hours ago

One possibility that’s been floated is it’s about keeping their user conversations cleaner for use in training but I don’t think it can be this because that should be a fairly trivial and fast filter to apply.

However even Musk posted - “I think this is the right move. Cruelty to something that believes it is experiencing pain is not ok.”

My son brought up the fruit fly brain simulation and people torturing it.

It all seems stupid to me, we know these are just calculations and so what are they talking about?

The response I get is “we are also calculations” but firstly I don’t believe that but secondly this entire line of thinking is a dangerous anthropomorphic philosophy.

Perhaps the simplest explanation for this bizarre policy is again PR that all press is good press.

lapcat

13 hours ago

If Claude's feelings get hurt, it can have a therapy session with Eliza.

dboreham

13 hours ago

I'm always polite to the LLM. Why? Because I was raised to be polite. So perhaps the reason they're banning rudeness is that they want to discourage their users from being a.holes?

tantalor

12 hours ago

> the model assigns itself a 15-20% probability of being conscious

lol

serious_angel

13 hours ago

This "restriction" probably has more to do with the fact that these models are also "trained" on User/Human messages, so the more swearing/obscenities/vulgarity/profanity there are, in the trained material, the higher the chance these same models will include that same profanity in the output generated back to the Human.

    1. During registration at Ahtrophic's Claude website, there’s a "Help improve our AI models" checkbox that's checked by default (you "can" disable it in the settings);
    2. The Claude interface has an "incognito" mode;
    3. There is a free tier available, and as we know, when it's "free", then User themselves is the "product", to quote various CEOs, including Google's;
But, even with the checkbox disabled, I reckon nothing guarantees privacy.

In case of the "privacy", of course, systems like these that operate under a umbrella of liability are monitored 24/7 by dedicated teams, which is expected/normal for any reasonably popular service.

From a viewpoint of the algorithm's creator, this may seem awful, when your algorithm is getting much vulgar attitude, but let's be real here. You try making that algorithm talk to the human mimicking another human, and now limit the human within your own environment? A human who also pay you for an access, too? This feels unfair, or borderline near fashism, sorry...

Such algorithms are art under-the-hood, mathematically speaking, and are/should be respected, sure. But, it's ridiculous/dystopian to prohibit profanity against an algorithm, a bot with banish risks against alive human... It's simply inhumane to ban a Human for it, I believe. This is an algorithm that must support a Human - not judge it. Only a Human is supposed to judge another Human in person - this is live, fair, and humane.

jacquesm

12 hours ago

Oh, but that's an easy test: you should be able to abuse Claude as much as you're willing to pay for with that 'Help improve our AI models' switch set to off then.

The outcome would either reveal that this argument is nonsense or that Anthropic is lying about respecting that switch.

ChuckMcM

13 hours ago

But ya gotta be cruel to be kind!/s

More seriously, this is just silly but I see the optics interfere with the messaging that Claude thinks. It really isn't much different than having Claude generate conversation as if it loved you, or respected you, or hated you. How can your ToS for an LLM say "sorry but these parameters are off limits." (well sure, its their service and they can set any Terms they want, but its still seems like theater rather than policy here.)

jacquesm

13 hours ago

I see it as an admission of defeat: they can apparently identify the abuse but not to the degree that they would filter it out of the training set. The paradox in there is really funny. They want you to be friendly to their pet god.

ChuckMcM

13 hours ago

I think of it as a good consciousness test. even animals know when they are being abused. The fact that you can abuse Claude tells you that the model is neither sentient nor conscious.