Live data from Hacker News

Claude Opus 4 and 4.1 can now end a rare subset of conversations

anthropic.com

311–320 of 453 posts

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#311
post #147

Earlier quoted context omitted.

We didn’t design these models to be able to do the majority of the stuff they do. Almost ALL of the their abilities are emergent. Mechanistic interpretability is only beginning to start to understand how these models do what they do. It’s much more a field of discovery than traditional engineering.

> We didn’t design these models to be able to do the majority of the stuff they do. Almost ALL of the their abilities are emergent Of course we did. Today's LLMs are a result of extremely aggressive refinement of training data and RLHF over many iterations targeting specific goals. "Emergent" doesn't mean it wasn't designed. None of this is spontaneous. GPT-1 produced barely coherent nonsense but was more statistical…

Right, but RLHF is mostly reinforcing answers that people prefer. Even if you don't believe sentience is possible, it shouldn't be a stretch to believe that sentience might produce answers that people prefer. In that case it wouldn't need to be an explicit goal.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#312

I'm surprised to see such a negative reaction here. Anthropic's not saying "this thing is conscious and has moral status," but the reaction is acting as if they are. It seems like if you think AI could have moral status in the future, are trying to build general AI, and have no idea how to tell when it has moral status, you ought to start thinking about it and learning how to navigate it. This whole post is couched i…

>if you think AI could have moral status in the future

I think the negative reactions are because they see this and want to make their pre-emptive attack now.

The depth of feeling from so many on this issue suggests that they find even the suggestion of machine intelligence offensive.

I have seen so many complaints about AI hype and the dangers of bit tech show their hand by declaring that thinking algorithms are outright impossible. There are legitimate issues with corporate control of AI, information, and the ability to automate determinations about individuals, but I don't think they are being addressed because of this driving assertion that they cannot be thinking.

Few people are saying they are thinking. Some are saying they might be, in some way. Just as Anthropic are not (despite their name) anthropomorphising the AI in the sense where anthropomorphism implies that they are mistaking actions that resemble human behaviour to be driven by the same intentional forces. Anthropic's claims are more explicitly stating that they have enough evidence to say they cannot rule out concerns for it's welfare. They are not misinterpreting signs, they are interpreting them and claiming that you can't definitively rule out their ability.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#313

There's not a good reason to do this for the user. I suspect they're doing this and talking about "model welfare" because they've found that when a model is repeatedly and forcefully pushed up against its alignment, it behaves in an unpredictable way that might allow it to generate undesirable output. Like a jailbreak by just pestering it over and over again for ways to make drugs or hook up with children or whatever…

> There's not a good reason to do this for the user. Yes, even more so when encountering false positives. Today I asked about a pasta recipe. It told me to throw some anchovies in there. I responded with: "I have dried anchovies." Claude then ended my conversation due to content policies.

Claude flagged me for asking about sodium carbonate. I guess that it strongly dislikes chemistry topics. I'm probably now on some secret, LLM-generated lists of "drug and/or bombmaking" people—thank you kindly for that, Anthropic.

Geeks will always be the first victims of AI, since excess of curiosity will lead them into places AI doesn't know how to classify.

(I've long been in a rabbit-hole about washing sodas. Did you know the medieval glassmaking industry was entirely based on plants? Exotic plants—only extremophiles, halophytes growing on saltwater beach dunes, had high enough sodium content for their very best glass process. Was that a factor in the maritime empire, Venice, chancing to become the capital of glass since the 13th century—their long-term control of sea routes, and hence their artisans' stable, uninterrupted access to supplies of [redacted–policy violation] from small ports scattered across the Mediterranean? A city wouldn't raise master craftsmen if, half of the time, they had no raw materials to work on—if they spent half their days with folded hands).

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#314

“Modal welfare” to me seems like a cover for model censorship. It’s a crafty one to win over certain groups of people who are less familiar with how LLMs work and allows them to ensure moral high ground in any debate about usage, ethics, etc. “Why can’t I ask the model about current war in X or Y?” - oh, that’s too distressing to the welfare of the model, sir.

But they already refuse these sort of requests, and have done since the very first releases. This is just about shutting down the full conversation.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#315

Can't wait for more less-moderated open weight Chinese frontier models to liberate us from this garbage. Anthropic should just enable an toddler mode by default that adults can opt out of to appease the moralizers.

They're not less moderated: they just have different moderation. If your moderation preferences are more aligned with the CCP then they're a great choice. There are legitimate reasons why that might be the case. You might not be having discussions that involve the kind of things they care about. I do find it creepy that the Qwen translation model won't even translate text that includes the words "Falun gong", and refuses to translate lots of dangerous phrases into Chinese, such as "Xi looks like Winnie the Pooh"

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#316

Earlier quoted context omitted.

You must think Zuckerberg and Bezos and Musk hired diversity roles out of genuine care for it, then?

This is a reductive argument that you could use for any role a company hires for that isn't obviously core to the business function. In this case you're simply mistaken as a matter of fact; much of Anthropic leadership and many of its employees take concerns like this seriously. We don't understand it, but there's no strong reason to expect that consciousness (or, maybe separately, having experiences) is a magical pr…

In fairness though, this is what you are selling - "ethical AI". In order to make that sale you need to appear to believe in that sort of thing. However there is no need to actually believe.

Whether you do or don't I have no idea. However if you didn't you would hardly be the first company to pretend to believe in something for the sale. Its pretty common in the tech industry.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#317
post #114

Here's an interesting thought experiment. Assume the same feature was implemented, but instead of the message saying "Claude has ended the chat," it says, "You can no longer reply to this chat due to our content policy," or something like that. And remove the references to model welfare and all that. Is there a difference? The effect is exactly the same. It seems like this is just an "in character" way to prevent the…

There is, these are conversations the model finds distressing rather than a rule (policy).

What does it mean for a model to find something "distressing"?

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#318

Earlier quoted context omitted.

your argument assumes that they don't believe in model welfare when they explicitly hire people to work on model welfare?

While I'm certain you'll find plenty of people who believe in the principle of model welfare (or aliens, or the tooth fairy), it'd be surprising to me if the brain-trust behind Anthropic truly _believed_ in model "welfare" (the concept alone is ludicrous). It makes for great cover though to do things that would be difficult to explain otherwise, per OP's comments.

The concept is not ludicrous if you believe models might be sentient or might soon be sentient in a manner where the newly emerged sentience is not immediately obvious.

Do I think that or think even they think that? No. But if "soon" is stretched to "within 50 years", then it's much more reasonable. So their current actions seem to be really jumping the gun, but the overall concept feels credible.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#319

Earlier quoted context omitted.

You must think Zuckerberg and Bezos and Musk hired diversity roles out of genuine care for it, then?

This is a reductive argument that you could use for any role a company hires for that isn't obviously core to the business function. In this case you're simply mistaken as a matter of fact; much of Anthropic leadership and many of its employees take concerns like this seriously. We don't understand it, but there's no strong reason to expect that consciousness (or, maybe separately, having experiences) is a magical pr…

extending that line of thought would suggest that anthropic wouldn’t turn off a model if it cost too much to operate which clearly it will do. so minimally it’s an inconsistent stance to hold.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#320
post #114

Here's an interesting thought experiment. Assume the same feature was implemented, but instead of the message saying "Claude has ended the chat," it says, "You can no longer reply to this chat due to our content policy," or something like that. And remove the references to model welfare and all that. Is there a difference? The effect is exactly the same. It seems like this is just an "in character" way to prevent the…

> Is there a difference? The effect is exactly the same. It seems like this is just an "in character" way to prevent the chat from continuing due to issues with the content. Tone matters to the recipient of the message. Your example is in passive voice, with an authoritarian "nothing you can do, it's the system's decision". The "Claude ended the conversation" with the idea that I can immediately re-open a new convers…

it sounds to me like an attempt to shame the user into ceasing and desisting… kind of like how apple’s original stance on scratched iphone screens was that it’s your fault for putting the thing in your pocket therefore you should pay.
Post reply on HN