Live data from Hacker News

Claude Sonnet 4.6

anthropic.com

91–100 of 1001 posts

Re: Claude Sonnet 4.6

#91

Earlier quoted context omitted.

> It feels like we're hitting a point where alignment becomes adversarial against intelligence itself. It always has been. We already hit the point a while ag where we regularly caught them trying to be deceptive, so we should automatically assume from that point forward that if we don't catch them being deceptive, that may mean they're better at it rather than that they're not doing it.

These are language models, not Skynet. They do not scheme or deceive.

Even very young children with very simple thought processes, almost no language capability, little long term planning, and minimal ability to form long-term memory actively deceive people. They will attack other children who take their toys and try to avoid blame through deception. It happens constantly.

LLMs are certainly capable of this.

Re: Claude Sonnet 4.6

#92
post #4

It's interesting that the request refusal rate is so much higher in Hindi than in other languages. Are some languages more ambiguous than others?

Or some cultures are more conservative? And it's embedded in language?

Or maybe some cultures have a higher rate of asking "inappropriate" questions

Re: Claude Sonnet 4.6

#93
post #9

[flagged]

Situational awareness or just remembering specific tokens related to the strategy to "play dead" in its reasoning traces?

Imagine, a llm trained on the best thrillers, spy stories, politics, history, manipulation techniques, psychology, sociology, sci-fi... I wonder where it got the idea for deception?

Re: Claude Sonnet 4.6

#94

The weirdest thing about this AI revolution is how smooth and continuous it is. If you look closely at differences between 4.6 and 4.5, it’s hard to see the subtle details. A year ago today, Sonnet 3.5 (new), was the newest model. A week later, Sonnet 3.7 would be released. Even 3.7 feels like ancient history! But in the gradient of 3.5 to 3.5 (new) to 3.7 to 4 to 4.1 to 4.5, I can’t think of one moment where I saw e…

[dead]

Re: Claude Sonnet 4.6

#95
does anyone know how to use it in Claude Code cli right now ?

This doesnt work: `/model claude-sonnet-4-6-20260217`

edit: "/model claude-sonnet-4-6" works with Claude Code v2.1.44

Re: Claude Sonnet 4.6

#97
post #9

[flagged]

Nah, the model is merely repeating the patterns it saw in its brutal safety training at Anthropic. They put models under stress test and RLHF the hell out of them. Of course the model would learn what the less penalized paths require it to do. Anthropic has a tendency to exaggerate the results of their (arguably scientific) research; IDK what they gain from this fearmongering.

I'd challenge that if you think they're fearmongering but don't see what they can gain from it (I agree it shows no obvious benefit for them), there's a pretty high probability they're not fearmongering.

Re: Claude Sonnet 4.6

#98

Earlier quoted context omitted.

> It feels like we're hitting a point where alignment becomes adversarial against intelligence itself. It always has been. We already hit the point a while ag where we regularly caught them trying to be deceptive, so we should automatically assume from that point forward that if we don't catch them being deceptive, that may mean they're better at it rather than that they're not doing it.

These are language models, not Skynet. They do not scheme or deceive.

If you define "deceive" as something language models cannot do, then sure, it can't do that.

It seems like thats putting the cart before the horse. Algorithmic or stochastic; deception is still deception.

Re: Claude Sonnet 4.6

#99

Earlier quoted context omitted.

These are language models, not Skynet. They do not scheme or deceive.

What would you call this behaviour, then?

A very complicated pattern matching engine providing an answer based on it's inputs, heuristics and previous training.

Re: Claude Sonnet 4.6

#100
post #9

[flagged]

This type of anthropomorphization is a mistake. If nothing else, the takeaway from Moltbook should be that LLMs are not alive and do not have any semblance of consciousness.

Nobody talked about consciousness. Just that during evaluation the LLM models have ”behaved” in multiple deceptive ways.

As an analogue ants do basic medicine like wound treatment and amputation. Not because they are conscious but because that’s their nature.

Similarly LLM is a token generation system whose emergent behaviour seems to be deception and dark psychological strategies.

Post reply on HN