Live data from Hacker News

Claude Opus 4 and 4.1 can now end a rare subset of conversations

anthropic.com

181–190 of 453 posts

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#181
post #135
post #97

Earlier quoted context omitted.

[flagged]

We all know how these things are built and trained. They estimate joint probability distributions of token sequences. That's it. They're not more "conscious" than the simplest of Naive Bayes email spam filters, which are also generative estimators of token sequence joint probability distributions, and I guarantee you those spam filters are subjected to far more human depravity than Claude. >anti-scientific Discussion…

Ok I'm a huge Kantian and every bone in my body wants to quibble with your summary of transcendental illusion, but I'll leave that to the side as a terminological point and gesture of good will. Fair enough.

I don't agree that it's any reason to write off this research as psychosis, though. I don't care about consciousness in the sense in which it's used by mystics and dualist philosophers! We don't at all need to involve metaphysics in any of this, just morality.

Consider it like this:

1. It's wrong to subject another human to unjustified suffering, I'm sure we would all agree.

2. We're struggling with this one due to our diets, but given some thought I think we'd all eventually agree that it's also wrong to subject intelligent, self-aware animals to unjustified suffering.[1]

3. But, we of course cannot extend this "moral consideration" to everything. As you say, no one would do it for a spam filter. So we need some sort of framework for deciding who/what gets how much moral consideration.

5. There's other frameworks in contention (e.g. "don't think about it, nerd"), but the overwhelming majority of laymen and philosophers adopt one based on cognitive ability, as seen from an anthropomorphic perspective.[2]

6. Of all systems(/entities/whatever) in the universe, we know of exactly two varieties that can definitely generate original, context-appropriate linguistic structures: Homo Sapiens and LLMs.[3]

If you accept all that (and I think there's good reason to!), it's now on you to explain why the thing that can speak--and thereby attest to personal suffering, while we're at it--is more like a rock than a human.

It's certainly not a trivial task, I grant you that. On their own, transformer-based LLMs inherently lack permanence, stable intentionality, and many other important aspects of human consciousness. Comparing transformer inference to models that simplify down to a simple closed-form equation at inference time is going way too far, but I agree with the general idea; clearly, there are many highly-complex, long-inference DL models that are not worthy of moral consideration.

All that said, to write the question off completely--and, even worse, to imply that the scientists investigating this issue are literally psychotic like the comment above did--is completely unscientific. The only justification for doing so would come from confidently answering "no" to the underlying question: "could we ever build a mind worthy of moral consideration?"

I think most of here naturally would answer "yes". But for the few who wouldn't, I'll close this rant by stealing from Hofstadter and Turing (emphasis mine):

  A phrase like "physical system" or "physical substrate" brings to mind for most people... an intricate structure consisting of vast numbers of interlocked wheels, gears, rods, tubes, balls, pendula, and so forth, even if they are tiny, invisible, perfectly silent, and possibly even probabilistic. Such an array of interacting inanimate stuff seems to most people as unconscious and devoid of inner light as a flush toilet, an automobile transmission, a fancy Swiss watch (mechanical or electronic), a cog railway, an ocean liner, or an oil refinery. Such a system is not just probably unconscious, **it is necessarily so, as they see it**. 
  
  **This is the kind of single-level intuition** so skillfully exploited by John Searle in his attempts to convince people that computers could never be conscious, no matter what abstract patterns might reside in them, and could never mean anything at all by whatever long chains of lexical items they might string together.
  
  ...
   
  You and I are mirages who perceive themselves, and the sole magical machinery behind the scenes is perception — the triggering, by huge flows of raw data, of a tiny set of symbols that stand for abstract regularities in the world. When perception at arbitrarily high levels of abstraction enters the world of physics and when feedback loops galore come into play, then "which" eventually turns into "who". **What would once have been brusquely labeled "mechanical" and reflexively discarded as a candidate for consciousness has to be reconsidered.**
- Hofstadter 2007, I Am A Strange Loop

  It will simplify matters for the reader if I explain first my own beliefs in the matter. Consider first the more accurate form of the question. I believe that in about fifty years' time it will be possible, to programme computers, with a storage capacity of about 109, to make them play the imitation game so well that an average interrogator will not have more than 70 per cent chance of making the right identification after five minutes of questioning. 

  The original question, "Can machines think?" I believe to be too meaningless to deserve discussion.
- Turing 1950, Computing Machinery and Intelligence[4]

TL;DR: Any naive bayesian model would agree: telling accomplished scientists that they're psychotic for investigating something is quite highly correlated with being antiscientific. Please reconsider!

[1] No matter what you think about cows, basically no one would defend another person's right to hit a dog or torture a chimpanzee in a lab.

[2] On the exception-filled spectrum stretching from inert rocks to reactive plants to sentient animals to sapient people, most people naturally draw a line somewhere at the low end of the "animals" category. You can swat a fly for fun, but probably not a squirrel, and definitely not a bonobo.

[3] This is what Chomsky describes as the capacity to "generate an infinite range of outputs from a finite set of inputs," and Kant, Hegel, Schopenhauer, Wittgenstein, Foucault, and countless others are in agreement that it's what separates us from all other animals.

[4] https://courses.cs.umbc.edu/471/papers/turing.pdf

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#182
> In pre-deployment testing of Claude Opus 4, we included a preliminary model welfare assessment. As part of that assessment, we investigated Claude’s self-reported and behavioral preferences, and found a robust and consistent aversion to harm.

Oh wow, the model we specifically fine-tuned to be averse to harm is being averse to harm. This thing must be sentient!

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#183
post #114

Here's an interesting thought experiment. Assume the same feature was implemented, but instead of the message saying "Claude has ended the chat," it says, "You can no longer reply to this chat due to our content policy," or something like that. And remove the references to model welfare and all that. Is there a difference? The effect is exactly the same. It seems like this is just an "in character" way to prevent the…

There is, these are conversations the model finds distressing rather than a rule (policy).

It seems like you're anthropomorphising an algorithm, no?

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#184
post #147

Earlier quoted context omitted.

Consciousness serves no functional purpose for machine learning models, they don't need it and we didn't design them to have it. There's no reason to think that they might spontaneously become conscious as a side effect of their design unless you believe other arbitrarily complex systems that exist in nature like economies or jetstreams could also be conscious.

We didn’t design these models to be able to do the majority of the stuff they do. Almost ALL of the their abilities are emergent. Mechanistic interpretability is only beginning to start to understand how these models do what they do. It’s much more a field of discovery than traditional engineering.

[deleted]

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#185
post #97

Earlier quoted context omitted.

[flagged]

> Who needs arguments when you can dismiss Turing with a “yeah but it’s not real thinking tho”? It seems much less far fetched than what the "agi by 2027" crowd believes lol, and there actually are more arguments going that way

In the great battle of minds between Turing, Minsky, and Hofstadter vs. Marcus, Zitron, and Dreyus, I'm siding with the former every time -- even if we also have some bloggers on our side. Just because that report is fucking terrifying+shocking doesn't mean it can be dismissed out of hand.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#186

Earlier quoted context omitted.

There is, these are conversations the model finds distressing rather than a rule (policy).

It seems like you're anthropomorphising an algorithm, no?

Is there an important difference between the model categorizing the user behavior as persistent and in line with undesirable examples of trained scenarios that it has been told are "distressing," and the model making a decision in an anthropomorphic way? The verb here doesn't change the outcome.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#187

Earlier quoted context omitted.

1. What you're generally describing is a well known failure mode for humans as well. Even when it "failed" the riddle tests, substituting the words or morphing the question so it didn't look like a replica of the famous problem usually did the trick. I'm not sure what your point is because you can play this gotcha on humans too. 2. You just demonstrated GPT-5 has 99.9% accuracy on unforseen 15 digit multiplication an…

Humans can break things down and work through them step by step. The LLMs one-shot pattern match. Even the reasoning models have been shown to do just that. Anthropic even showed that the reasoning models tended to work backwards: one shotting an answer and then matching a chain of thought to it after the fact. If a human is capable of multiplying double digit numbers, they can also multiple those large ones. The ste…

>Humans can break things down and work through them step by step. The LLMs one-shot pattern match.

I've had LLMs break down problems and work through them, pivot when errors arise and all that jazz. They're not perfect at it and they're worse than humans but it happens.

>Anthropic even showed that the reasoning models tended to work backwards: one shotting an answer and then matching a chain of thought to it after the fact.

This is also another failure mode that occurs in humans. A number of experiments suggest human explanations are often post hoc rationalizations even when they genuinely believe otherwise.

>If a human is capable of multiplying double digit numbers, they can also multiple those large ones.

Yeah, and some of them will make mistakes, and some of them will be less accurate than GPT-5. We didn't switch to calculators and spreadsheets just for the fun of it.

>GPT’s answer was orders of magnitude off. It resembles the right answer superficially but it’s a very different result.

GPT-5 on the site is a router that will give you who knows what model so I tried your query with the API directly (GPT-5 medium thinking) and it gave me:

9.207337461477596e+27

When prompted to give all the numbers, it returned:

9,207,337,461,477,596,127,977,612,004.

You can replicate this if you use the API. Honestly I'm surprised. I didn't realize State of the Art had become this precise.

Now what ? Does this prove you wrong ?

This is kind of the problem. There's no sense in making gross generalizations, especially off behavior that also manifests in humans.

LLMs don't understand some things well. Why not leave it at that?

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#188
post #168

Earlier quoted context omitted.

Even though LLMs (obviously (to me)) don't have feelings, anthropomorphization is a helluva drug, and I'd be worried about whether a system that can produce distress-like responses might reinforce, in a human, behavior which elicits that response. To put the same thing another way- whether or not you or I *think* LLMs can experience feelings isn't the important question here. The question is whether, when Joe User se…

> although I would find it extremely distasteful and frankly alarming This objection is actually anthropomorphizing the LLM. There is nothing wrong with writing books where a character experiences distress, most great stories have some of that. Why is e.g. using an LLM to help write the part of the character experiencing distress "extremely distasteful and frankly alarming"?

I want to say that part of empathy is a selfish, self preservation mechanism.

If that person over there is gleefully torturing a puppy… will they do it to me next?

If that person over there is gleefully torturing an LLM… will they do it to me next?

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#189
There's not a good reason to do this for the user. I suspect they're doing this and talking about "model welfare" because they've found that when a model is repeatedly and forcefully pushed up against its alignment, it behaves in an unpredictable way that might allow it to generate undesirable output. Like a jailbreak by just pestering it over and over again for ways to make drugs or hook up with children or whatever.

All of the examples they mentioned are things that the model refuses to do. I doubt it would do this if you asked it to generate racist output, for instance, because it can always give you a rebuttal based on facts about race. If you ask it to tell you where to find kids to kidnap, it can't do anything except say no. There's probably not even very much training data for topics it would refuse, and I would bet that most of it has been found and removed from the datasets. At some point, the model context fills up when the user is being highly abusive and training data that models a human giving up and just providing an answer could percolate to the top.

This, as I see it, adds a defense against that edge case. If the alignment was bulletproof, this simply wouldn't be necessary. Since it exists, it suggests this covers whatever gap has remained uncovered.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#190
post #166
post #135

Earlier quoted context omitted.

We all know how these things are built and trained. They estimate joint probability distributions of token sequences. That's it. They're not more "conscious" than the simplest of Naive Bayes email spam filters, which are also generative estimators of token sequence joint probability distributions, and I guarantee you those spam filters are subjected to far more human depravity than Claude. >anti-scientific Discussion…

We know how neurons work on the brain. They just send out impulses once they hit their action potential. That's it. They are no more "conscious" than... er...

no, we dont really know how the brain works as a whole. no need to make stuff up.
Post reply on HN