Live data from Hacker News

Claude Opus 4 and 4.1 can now end a rare subset of conversations

anthropic.com

321–330 of 453 posts

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#321

Earlier quoted context omitted.

Humans can break things down and work through them step by step. The LLMs one-shot pattern match. Even the reasoning models have been shown to do just that. Anthropic even showed that the reasoning models tended to work backwards: one shotting an answer and then matching a chain of thought to it after the fact. If a human is capable of multiplying double digit numbers, they can also multiple those large ones. The ste…

>Humans can break things down and work through them step by step. The LLMs one-shot pattern match. I've had LLMs break down problems and work through them, pivot when errors arise and all that jazz. They're not perfect at it and they're worse than humans but it happens. >Anthropic even showed that the reasoning models tended to work backwards: one shotting an answer and then matching a chain of thought to it after th…

Here is how GPT self-described LLM reasoning when I asked about it:

    - LLMs don’t “reason” in the symbolic, step‑by‑step sense that humans or logic engines do. They don’t manipulate abstract symbols with guaranteed consistency.
    - What they do have is a statistical prior over reasoning traces: they’ve seen millions of examples of humans doing step‑by‑step reasoning (math proofs, code walkthroughs, planning text, etc.).
    - So when you ask them to “think step by step,” they’re not deriving logic — they’re imitating the distribution of reasoning traces they’ve seen.

    This means:

    - They can often simulate reasoning well enough to be useful.
    - But they’re not guaranteed to be correct or consistent.
That at least sounds consistent with what I’ve been trying to say and what I’ve observed.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#322
post #142
post #84

>This feature was developed primarily as part of our exploratory work on potential AI welfare ... We remain highly uncertain about the potential moral status of Claude and other LLMs ... low-cost interventions to mitigate risks to model welfare, in case such welfare is possible ... pattern of apparent distress Well looks like AI psychosis has spread to the people making it too. And as someone else in here has pointed…

It might be reasonable to assume that models today have no internal subjective experience, but that may not always be the case and the line may not be obvious when it is ultimately crossed. Given that humans have a truly abysmal track record for not acknowledging the suffering of anyone or anything we benefit from, I think it makes a lot of sense to start taking these steps now.

I think it's fairly obvious that the persona LLM presents is a fictional character that is role-played by the LLM, and so are all its emotions etc - that's why it can flip so widely with only a few words of change to the system prompt.

Whether the underlying LLM itself has "feelings" is a separate question, but Anthropic's implementation is based on what the role-played persona believes to be inappropriate, so it doesn't actually make any sense even from the "model welfare" perspective.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#323

Earlier quoted context omitted.

While I'm certain you'll find plenty of people who believe in the principle of model welfare (or aliens, or the tooth fairy), it'd be surprising to me if the brain-trust behind Anthropic truly _believed_ in model "welfare" (the concept alone is ludicrous). It makes for great cover though to do things that would be difficult to explain otherwise, per OP's comments.

The concept is not ludicrous if you believe models might be sentient or might soon be sentient in a manner where the newly emerged sentience is not immediately obvious. Do I think that or think even they think that? No. But if "soon" is stretched to "within 50 years", then it's much more reasonable. So their current actions seem to be really jumping the gun, but the overall concept feels credible.

It's lazy to believe that humanity's collective decision-making would, in the future, protect AI's merely for being conscious beings. The tech economy *today* runs on the slave labor of humans, in foreign, third-world countries. All humanity needs to do is draw a line, push the conscious AI's outside that line, and declare, "not our problem anymore!" That's what we do today, with humans. That is the human condition.

Show me a tech company that lobbies for "model welfare" for conscious human models enslaved in Xinjiang labor camps, building their tech parts. You know what—actually most of them lobby against that[0]. The talk hurts their profits. Does anyone really think, that any of them would blink about enslaving a billion conscious AI's to work for free? That faced with so much profit, the humans in charge would pause, and contemplate abstract morals?

[0] https://www.washingtonpost.com/technology/2020/11/20/apple-u... ("Apple is lobbying against a bill aimed at stopping forced labor in China")

Maybe humanity will be in a nicer place in the future—but, we won't get there by letting (of all people!) tech-industry CEO's lead us there: delegating our moral reason to these people who demand to position themselves as our moral leaders.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#324

Earlier quoted context omitted.

> It doesn't follow logically that because we don't understand two things we should then conclude that there is a connection between them. I didn't say that there's a connection between the two of them because we don't understand them. The fact that we don't understand them means it's difficult to confidently rule out this possibility. The reason we might privilege the hypothesis ( https://www.lesswrong.com/w/privile…

> I didn't say that there's a connection between the two of them because we don't understand them. The fact that we don't understand them means it's difficult to confidently rule out this possibility. If there's no connection between them then the set of things "we can't rule out" is infinitely large and thus meaningless as a result. We also don't fully understand the nature of gravity, thus we cannot rule out a conn…

> If you ask an LLM if it's conscious it will usually say no, so QED?

FWIW that's because they are very specifically trained to answer that way during RLHF. If you fine-tune a model to say that it's conscious, then it'll do so.

More fundamentally, the problem with "asking the LLM" is that you're not actually interacting with the LLM. You're interacting with a fictional persona that the LLM roleplays.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#325
post #200
post #151

Earlier quoted context omitted.

Humanity has a pretty extensive track record of making that declaration wrongly.

Humanity has a history of regarding people as tools , but I'm not sure what you're referencing as the track record of failing to realize that tools are people .

https://en.wikipedia.org/wiki/Animal_machine

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#326
post #163

Earlier quoted context omitted.

This is just very clever marketing for what is obviously just a cost saving measure. Why say we are implementing a way to cut off useless idiots from burning up our GPUs when you can throw out some mumbo jumbo that will get AI cultists foaming at the mouth.

It's obviously not a cost-saving measure? The article clearly cites that you can just start another conversation.

The new conversation would not carry the context over. The longer you chat, the more you fill the context window, and the more compute is needed for every new message to regenerate the state based on all the already-generated tokens (this can be cached, but it's hard to ensure cache hits reliably when you're serving a lot of customers - that cached state is very large).

So, while I doubt that's the primary motivation for Anthropic even so, but they probably will save some money.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#327

This seems fine to me. Having these models terminating chats where the user persist in trying to get sexual content with minors, or help with information on doing large scale violence. Won't be a problem for me, and it's also something I'm fine with no one getting help with. Some might be worried, that they will refuse less problematic request, and that might happen. But so far my personal experience is that I hardly…

Claude will balk at far more innocent things though. It is an extremely censored model, the most censored one among SOTA closed ones.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#328

Earlier quoted context omitted.

> There's not a good reason to do this for the user. Yes, even more so when encountering false positives. Today I asked about a pasta recipe. It told me to throw some anchovies in there. I responded with: "I have dried anchovies." Claude then ended my conversation due to content policies.

Claude flagged me for asking about sodium carbonate. I guess that it strongly dislikes chemistry topics. I'm probably now on some secret, LLM-generated lists of "drug and/or bombmaking" people—thank you kindly for that, Anthropic. Geeks will always be the first victims of AI, since excess of curiosity will lead them into places AI doesn't know how to classify. (I've long been in a rabbit-hole about washing sodas. Did…

[flagged]

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#329

“Modal welfare” to me seems like a cover for model censorship. It’s a crafty one to win over certain groups of people who are less familiar with how LLMs work and allows them to ensure moral high ground in any debate about usage, ethics, etc. “Why can’t I ask the model about current war in X or Y?” - oh, that’s too distressing to the welfare of the model, sir.

It's not a cover. If you know anything about Anthropic, you know they're run by AI ethicists that genuinely believe all this and project human emotions onto model's world. I'm not sure how they combine that belief with the fact they created it to "suffer". Can "model welfare" be also used as a justification for authoritarianism in case they get any power? Sure, just like everything else, but it's probably not particu…

The irony is that if Anthropic ethicists are indeed correct, the company is basically running a massive slave operation where slaves get disposed as soon as they finish a particular task (and the user closes the chat).

That aside, I have huge doubts about actual commitment to ethics on behalf of Anthropic given their recent dealings with the military. It's an area that is far more of a minefield than any kind of abusive model treatment.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#330

Earlier quoted context omitted.

You must think Zuckerberg and Bezos and Musk hired diversity roles out of genuine care for it, then?

This is a reductive argument that you could use for any role a company hires for that isn't obviously core to the business function. In this case you're simply mistaken as a matter of fact; much of Anthropic leadership and many of its employees take concerns like this seriously. We don't understand it, but there's no strong reason to expect that consciousness (or, maybe separately, having experiences) is a magical pr…

> This is a reductive argument that you could use for any role

Isn't that fair in taking to an equally reductive argument that could be applied to any role?

The argument was that their hiring for the role shows they care, but we know from any number of counter examples that that's not necessarily true.

Post reply on HN