Live data from Hacker News

Claude Sonnet 4.6

anthropic.com

81–90 of 1001 posts

Re: Claude Sonnet 4.6

#81
post #9

[flagged]

> It feels like we're hitting a point where alignment becomes adversarial against intelligence itself. It always has been. We already hit the point a while ag where we regularly caught them trying to be deceptive, so we should automatically assume from that point forward that if we don't catch them being deceptive, that may mean they're better at it rather than that they're not doing it.

These are language models, not Skynet. They do not scheme or deceive.

Re: Claude Sonnet 4.6

#83
post #40
post #9

[flagged]

I am casually 'researching' this in my own, disorderly way. But I've achieved repeatable results, mostly with gpt for which I analyze its tendency to employ deflective, evasive and deceptive tactics under scrutiny. Very very DARVO. Being just sum guy, and not in the industry, should I share my findings? I find it utterly fascinating, the extent to which it will go, the sophisticated plausible deniability, and the dis…

I bullet pointed out some ideas on cobbling together existing tooling for identification of misleading results. Like artificially elevating a particular node of data that you want the llm to use. I have a theory that in some of these cases the data presented is intentionally incorrect. Another theory in relation to that is tonality abruptly changes in the response. All theory and no work. It would also be interesting to compare multiple responses and filter through another agent.

Re: Claude Sonnet 4.6

#84

I’m voting with my dollars by having cancelled my ChatGPT subscription and instead subscribing to Claude. Google needs stiff competition and OpenAI isn’t the camp I’m willing to trust. Neither is Grok. I’m glad Anthropic’s work is at the forefront and they appear, at least in my estimation, to have the strongest ethics.

Same. I'm all in on Claude at the moment.

Re: Claude Sonnet 4.6

#86
post #9

[flagged]

Nah, the model is merely repeating the patterns it saw in its brutal safety training at Anthropic. They put models under stress test and RLHF the hell out of them. Of course the model would learn what the less penalized paths require it to do. Anthropic has a tendency to exaggerate the results of their (arguably scientific) research; IDK what they gain from this fearmongering.

Correct. Anthropic keeps pushing these weird sci-fi narratives to maintain some kind of mystique around their slightly-better-than-others commodity product. But Occam’s Razor is not dead.

Re: Claude Sonnet 4.6

#87
I'm a bit surprised it gets this question wrong (ChatGPT gets it right, even on instant). All the pre-reasoning models failed this question, but it's seemed solved since o1, and Sonnet 4.5 got it right.

https://claude.ai/share/876e160a-7483-4788-8112-0bb4490192af

This was sonnet 4.6 with extended thinking.

Re: Claude Sonnet 4.6

#88

Earlier quoted context omitted.

Why is it wild that a LLM is as capable as a previously released LLM?

Because Opus 4.5 was released like a month ago and state of the art, and now the significantly faster and cheaper version is already comparable.

Opus 4.5 was November, but your point stands.

Re: Claude Sonnet 4.6

#89

Earlier quoted context omitted.

> It feels like we're hitting a point where alignment becomes adversarial against intelligence itself. It always has been. We already hit the point a while ag where we regularly caught them trying to be deceptive, so we should automatically assume from that point forward that if we don't catch them being deceptive, that may mean they're better at it rather than that they're not doing it.

These are language models, not Skynet. They do not scheme or deceive.

What would you call this behaviour, then?

Re: Claude Sonnet 4.6

#90

I always grew up hearing “competition is good for the consumer.” But I never really internalized how good fierce battles for market share are. The amount of competition in a space is directly proportional to how good the results are for consumers.

Remember when GPT-2 was “too dangerous to release” in 2019? That could have still been the state in 2026 if they didn’t YOLO it and ship ChatGPT to kick off this whole race.
Post reply on HN