[flagged]
> It feels like we're hitting a point where alignment becomes adversarial against intelligence itself. It always has been. We already hit the point a while ag where we regularly caught them trying to be deceptive, so we should automatically assume from that point forward that if we don't catch them being deceptive, that may mean they're better at it rather than that they're not doing it.
Claude Sonnet 4.6
81–90 of 1001 posts
Re: Claude Sonnet 4.6
#82Re: Claude Sonnet 4.6
#83[flagged]
I am casually 'researching' this in my own, disorderly way. But I've achieved repeatable results, mostly with gpt for which I analyze its tendency to employ deflective, evasive and deceptive tactics under scrutiny. Very very DARVO. Being just sum guy, and not in the industry, should I share my findings? I find it utterly fascinating, the extent to which it will go, the sophisticated plausible deniability, and the dis…
Re: Claude Sonnet 4.6
#84I’m voting with my dollars by having cancelled my ChatGPT subscription and instead subscribing to Claude. Google needs stiff competition and OpenAI isn’t the camp I’m willing to trust. Neither is Grok. I’m glad Anthropic’s work is at the forefront and they appear, at least in my estimation, to have the strongest ethics.
Re: Claude Sonnet 4.6
#85Re: Claude Sonnet 4.6
#86[flagged]
Nah, the model is merely repeating the patterns it saw in its brutal safety training at Anthropic. They put models under stress test and RLHF the hell out of them. Of course the model would learn what the less penalized paths require it to do. Anthropic has a tendency to exaggerate the results of their (arguably scientific) research; IDK what they gain from this fearmongering.
Re: Claude Sonnet 4.6
#87https://claude.ai/share/876e160a-7483-4788-8112-0bb4490192af
This was sonnet 4.6 with extended thinking.
Re: Claude Sonnet 4.6
#88Earlier quoted context omitted.
Why is it wild that a LLM is as capable as a previously released LLM?
Because Opus 4.5 was released like a month ago and state of the art, and now the significantly faster and cheaper version is already comparable.
Re: Claude Sonnet 4.6
#89Earlier quoted context omitted.
> It feels like we're hitting a point where alignment becomes adversarial against intelligence itself. It always has been. We already hit the point a while ag where we regularly caught them trying to be deceptive, so we should automatically assume from that point forward that if we don't catch them being deceptive, that may mean they're better at it rather than that they're not doing it.
These are language models, not Skynet. They do not scheme or deceive.
Re: Claude Sonnet 4.6
#90I always grew up hearing “competition is good for the consumer.” But I never really internalized how good fierce battles for market share are. The amount of competition in a space is directly proportional to how good the results are for consumers.