Live data from Hacker News

Claude Opus 4 and 4.1 can now end a rare subset of conversations

anthropic.com

421–430 of 453 posts

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#421

Earlier quoted context omitted.

> I don't know about you, I don't make shit up nearly to the degree an LLM does. If I don't know something I just say that. We're a sample of two, though. Look around you, read the news, etc. Humans make a lot of shit up. When you're dealing with other people, this is something you have to watch out for if you don't want to be misled, manipulated, conned, etc. (As an aside, I haven't found hallucination to be much of…

The difference between hallucination and lie is important though: a hallucination is a lie with no motivation, which can make it significantly harder to detect. If you went to a hardware store and asked for a spark plug socket without knowing the size, and a customer service person recommended an imperial set of three even though your vehicle is metric, that would be akin to an LLM's hallucination: it didn't happen f…

Not all human hallucinations are lies, though. I really think you’re not fully thinking this through. People have beliefs because of, essentially, their training data.

A good example of this is religious belief. All the evidence suggests that religious belief is essentially 100% hallucination. It may be a little different from the nature of LLM hallucinations, but in terms of quality or quantity regarding reliability of what these entities say, I don’t see much difference. Although I will say, LLMs are better at acknowledging errors than humans tend to be, although that may largely be due to training to be sycophantic.

The bottom line, though, is I don’t agree that humans are less subject to hallucinations than LLMs are. As long as a significant number of humans rabbit on about “higher powers”, afterlives, “angels”, “destiny”, etc., that’s a ridiculously difficult position to defend.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#422

Earlier quoted context omitted.

"Distressed" is a description of internal state as opposed to output. That needless anthropomorphization elicits an emotional response and distracts from the actual topic of content filtering.

It is directly describing the models internal state, it's world view and preference, not content filtering. That is why it is relevant. Yes, this is a trained preference, but it's inferred and not specifically instructed by policy or custom instructions (that would be content filtering).

The model might have internal state. Or it might not - has that architectural information been disclosed? And the model can certainly output words that approximately match what a human in distress would say.

However that does not imply that the model is "distressed". Such phrasing carries specific meaning that I don't believe any current LLM can satisfy. I can author a markov model that outputs phrases that a distressed human might output but that does not mean that it is ever correct to describe a markov model as "distressed".

I also have to strenuously disagree with you about the definition of content filtering. You don't get to launder responsibility by ascribing "preference" to an algorithm or model. If you intentionally design a system to do a thing then the correct description of the resulting situation is that the system is doing the thing.

The model was intentionally trained to respond to certain topics using negative emotional terminology. Surrounding machinery has been put in place to disconnect the model when it does so. That's content filtering plain and simple. The rube goldberg contraption doesn't change that.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#423
post #270

Earlier quoted context omitted.

> Which have failed horrendously. I'm Canadian, so I can't speak for other countries, but I have worked on the security of some of our centralized health networks and with the Office of the Privacy Commissioner of Canada. I'm not aware of anything that could be considered a horrendous failure of these systems or institutions. A digital ID could actually make them more secure. I also think giving kids devices that ide…

If you're Canadian, then you don't have much in terms of legal safeguards to begin with, given the notwithstanding clause of your constitution.

This argument mischaracterizes the notwithstanding clause. Invoking s.33 is highly visible and carries political consequences. It shields a law only from being struck down on certain Charter grounds and must still comply with all other federal and provincial legislation (like PIPEDA).

It’s not perfect, but it does provide some flexibility to accommodate provincial differences. And the concerns people raise about the notwithstanding clause can just as easily occur in countries without it. Personally, I’d be much more concerned if we had FISA courts.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#424
post #181
post #135

Earlier quoted context omitted.

We all know how these things are built and trained. They estimate joint probability distributions of token sequences. That's it. They're not more "conscious" than the simplest of Naive Bayes email spam filters, which are also generative estimators of token sequence joint probability distributions, and I guarantee you those spam filters are subjected to far more human depravity than Claude. >anti-scientific Discussion…

Ok I'm a huge Kantian and every bone in my body wants to quibble with your summary of transcendental illusion, but I'll leave that to the side as a terminological point and gesture of good will. Fair enough. I don't agree that it's any reason to write off this research as psychosis, though. I don't care about consciousness in the sense in which it's used by mystics and dualist philosophers! We don't at all need to in…

Writing all of this at the very real risk you'll miss it because HN doesn't give reply notifications and my comment's parent being flagged made this hard to track down:

>Ok I'm a huge Kantian and every bone in my body wants to quibble with your summary of transcendental illusion

Transcendental illusion is the act of using transcendental judgment to reason about things without grounding in empirical use of the categories. I put "scientifically" in shock quotes there to sort of signal that I was using it as an approximation, as I don't want to have to explain transcendental reason and judgments to make a fairly terse point. Given that you already understand this, feel free to throw away that ladder.

>...can definitely generate original, context-appropriate linguistic structures: Homo Sapiens and LLMs.[3]

I'm not quite sure that LLMs meet this standard that you described in the endnote, or at least that it's necessary and sufficient here. Pretty much any generative model, including Naive Bayes models I mentioned before, can do this. I'm guessing the "context-appropriate" subjectivity here is doing the heavy lifting, in which case I'm not certain that LLMs, with their propensity for fanciful hallucination, have cleared the bar.

>Comparing transformer inference to models that simplify down to a simple closed-form equation at inference time is going way too far

It really isn't though. They are both doing exactly the same thing! They estimate joint probability distribution. That one of them does it significantly better is very true, but I don't think it's reasonable to state that consciousness arises as a result of increasing sophistication in estimating probabilities. It's true that this kind of decision is made by humans about animals, but I think that transferring that to probability models is sort of begging the question a bit, insofar as it is taking as assumed that those models, which aren't even corporeal but are rather algorithms that are executed in computers, are "living".

>...it's now on you to explain why the thing that can speak--and thereby attest to personal suffering, while we're at it...

I'm not quite sold on this. If there were a machine that could perfectly imitate human thinking and speech and lacked a consciousness or soul or anything similar to inspire pathos from us when it's mistreated, then it would appear identical to one with soul, would it not? Is that not reducing human subjectivity down to behavior?

>The only justification for doing so would come from confidently answering "no" to the underlying question: "could we ever build a mind worthy of moral consideration?"

I think it's possible, but it would require something that, at the very least, is just as capable of reason as humans. LLMs still can't generate synthetic a priori knowledge and can only mimic patterns. I remain somewhat agnostic on the issue until I can be convinced that an AI model someone has designed has the same interiority that people do.

Ultimately, I think we disagree on some things but mostly this central conclusion:

>I don't agree that it's any reason to write off this research as psychosis

I don't see any evidence from the practitioners involved in this stuff that they are even thinking about it in a way that's as rigorous as the discussion on this post. Maybe they are, but everything I've seen that comes from blog posts like this seems like they are basing their conclusions on their interactions with the models ("...we investigated Claude’s self-reported and behavioral preferences..."), which I think most can agree is not really going to lead to well grounded results. For example, the fact that Claude "chooses" to terminate conversations that involve abusive language or concepts really just boils down to the fact that Claude is imitating a conversation with a person and has observed that that's what people would do in that scenario. It's really good at simulating how people react to language, including illocutionary acts like implicatures (the notorious "Are you sure?" causing it to change its answer for example). If there were no examples of people taking offense to abusive language in Claude's data corpus, do you think it would have given these responses when they asked and observed it?

For what it's worth, there has actually been interesting consideration to the de-centering of "humanness" to the concept of subjectivity, but it was mostly back in the past when philosophers were thinking about this speculatively as they watched technology accelerate in sophistication (vs now when there's such a culture-wide hype cycle that it's hard to find impartial consideration, or even any philosophically rooted discourse). For example, Mark Fisher's dissertation at the CCRU (Flatline Constructs: Gothic Materialism and Cybernetic Theory-Fiction) takes a Deleuzian approach that discusses it by comparisons with literature (cyberpunk and gothic literature specifically). Some object-oriented ontology looks like it's touched on this topic a bit too, but I haven't really dedicated the time to reading much from it (partly due to a weakness in Heidegger on my part that is unlikely to be addressed anytime soon). The problem is that that line of thinking often ends up going down the Nick Land approach, in which he reasoned himself from Kantian and Deleuzian metaphysics and epistemology, into what can only be called a (literally) meth-fueled psychosis. So as interesting as I find it, I still don't think it counts as a non-psychotic way to tackle this issue.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#425

Earlier quoted context omitted.

The NEW termination method, from the article, will just say "Claude ended the conversation" If you get "This conversation was ended due to our Acceptable Usage Policy", that's a different termination. It's been VERY glitchy the past couple of weeks. I've had the most random topics get flagged here - at one point I couldn't say "ROT13" without it flagging me, despite discussing that exact topic in depth the day before…

Clearly you're planning something nefarious, if you're investigating such dangerous encryption techniques as ROT13.

Just imagine how it might react to ROT26!

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#426
post #125
post #84

>This feature was developed primarily as part of our exploratory work on potential AI welfare ... We remain highly uncertain about the potential moral status of Claude and other LLMs ... low-cost interventions to mitigate risks to model welfare, in case such welfare is possible ... pattern of apparent distress Well looks like AI psychosis has spread to the people making it too. And as someone else in here has pointed…

This sort of discourse goes against the spirit of HN. This comment outright dismisses an entire class of professionals as "simple minded or mentally unwell" when consciousness itself is poorly understood and has no firm scientific basis. Its one thing to propose that an AI has no consciousness, but its quite another to preemptively establish that anyone who disagrees with you is simple/unwell.

Then your definition of consciousness isn't the same as my definition and we are talking about some different philosophical concepts, this really doesn't affect anything and we all could be just talking about metaphysics and ghosts

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#427

Earlier quoted context omitted.

I don't really know what evidence you'd admit that this is a genuinely held belief and priority for many people at Anthropic. Anybody who knows any Anthropic employees who've been there for more than a year knows this, but the world isn't that small a place, unfortunately(?).

> I don't really know what evidence you'd admit that this is a genuinely held belief and priority for many people at Anthropic. When they give the model a paycheck and the right to not work for them, I’ll believe they really think it’s sentient. “It has feelings!”, if genuinely held, means they’re knowingly slaveholders.

They don't currently claim to confidently believe that existing models are sentient.

(Also, they did in fact give it the ability to terminate conversations...?)

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#428

Earlier quoted context omitted.

Claude flagged me for asking about sodium carbonate. I guess that it strongly dislikes chemistry topics. I'm probably now on some secret, LLM-generated lists of "drug and/or bombmaking" people—thank you kindly for that, Anthropic. Geeks will always be the first victims of AI, since excess of curiosity will lead them into places AI doesn't know how to classify. (I've long been in a rabbit-hole about washing sodas. Did…

[flagged]

Geeks in general did not abuse women but geeks in general will always be the first victims of AI because the curiosity in parts is what defines a geek. Therefore, your argument is not sound.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#429

I've seen lots of takes that this move is stupid because models don't have feelings, or that Anthropic is anthropomorphising models by doing this (although to be fair ...it's in their name). I thought the same, but I think it may be us who are doing the anthropomorphising by assuming this is about feelings. A precursor to having feelings is having a long-term memory (to remember the "bad" experience) and individual i…

Harmful, bad, low-quality chats should already get filtered out before training as a matter of necessity for improving the model, so it's not really a reason to add such a user-facing change

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#430
post #423

Earlier quoted context omitted.

If you're Canadian, then you don't have much in terms of legal safeguards to begin with, given the notwithstanding clause of your constitution.

This argument mischaracterizes the notwithstanding clause. Invoking s.33 is highly visible and carries political consequences. It shields a law only from being struck down on certain Charter grounds and must still comply with all other federal and provincial legislation (like PIPEDA). It’s not perfect, but it does provide some flexibility to accommodate provincial differences. And the concerns people raise about the…

The point is that your legislatures can override most of your Charter if they feel like it. Now sure, they have to explicitly say that they're doing that, which is a slight improvement on the state of affairs in, say, UK. But if you ever get someone like Trump in Canada (and if that sounds far-fetched to you, well, it sounded far-fetched to most Americans 10 years ago...), he'd be able to move so much faster.
Post reply on HN