> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. And Opus 4.8 is still cheaper for a higher pass rate (much less open weight models like GLM 5.2) so not sure why I'd use Sonnet except on the…
They spent months hyping up Mythos and ended up with it banned. I’d assume they want to both differentiate their products and appeal to regulators here
Claude Sonnet 5
161–170 of 822 posts
Re: Claude Sonnet 5
#162Earlier quoted context omitted.
There's two classes of models now - the cybersecurity ones that none of us are getting, and the 'safe' models released for general consumption. This is letting us know which side of the divide it sits on.
There's also Chinese models, which aren't trying to self-limit capabilities.
Also, I wouldn’t expect Mythos-class models to be allowed to be openly released by the CCP. Thinking otherwise is pure naivety.
Re: Claude Sonnet 5
#163Earlier quoted context omitted.
There's two classes of models now - the cybersecurity ones that none of us are getting, and the 'safe' models released for general consumption. This is letting us know which side of the divide it sits on.
There's also Chinese models, which aren't trying to self-limit capabilities.
Re: Claude Sonnet 5
#164Re: Claude Sonnet 5
#165Earlier quoted context omitted.
Why do you think they are bragging? Anthropic has long been the company to give us by far the most in-depth information about their models, both positive and negative. I read this as them just stating a fact about this model that users would want to know.
The preceding sentence is >Our safety assessments found that Sonnet 5 shows an overall lower rate of undesirable behaviors than Sonnet 4.6, and is generally safer to use in agentic contexts. which is obviously painting that as a good thing. So reading the next sentence as "in other good news" is reasonable.
Re: Claude Sonnet 5
#166Re: Claude Sonnet 5
#167Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models. I have been using Sonnet 4.6 more than Opus, because I'm mostly doing agent-assisted development and not fully agent-driven development. This announcement does not make me positive, I have fou…
From my own experience, GLM-5.2 generally cost more tokens and much more slow.
Re: Claude Sonnet 5
#168Unfortunately that means I won't be using it at work for now.
Re: Claude Sonnet 5
#169Re: Claude Sonnet 5
#170It seems being incompetent is a feature now...