Live data from Hacker News

Claude Sonnet 5

anthropic.com

161–170 of 822 posts

Re: Claude Sonnet 5

#161
post #42

> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. And Opus 4.8 is still cheaper for a higher pass rate (much less open weight models like GLM 5.2) so not sure why I'd use Sonnet except on the…

They spent months hyping up Mythos and ended up with it banned. I’d assume they want to both differentiate their products and appeal to regulators here

I'm starting to think it discovered a 0-day held hidden by our government.

Re: Claude Sonnet 5

#162
post #89

Earlier quoted context omitted.

There's two classes of models now - the cybersecurity ones that none of us are getting, and the 'safe' models released for general consumption. This is letting us know which side of the divide it sits on.

There's also Chinese models, which aren't trying to self-limit capabilities.

…as long as you don’t ask them about certain dates or squares.

Also, I wouldn’t expect Mythos-class models to be allowed to be openly released by the CCP. Thinking otherwise is pure naivety.

Re: Claude Sonnet 5

#163
post #89

Earlier quoted context omitted.

There's two classes of models now - the cybersecurity ones that none of us are getting, and the 'safe' models released for general consumption. This is letting us know which side of the divide it sits on.

There's also Chinese models, which aren't trying to self-limit capabilities.

Surely the Chinese government will see US gov's intervention and say "Government control of business is stupid, our industry will have more independence from CCP control for the benefit of the world".

Re: Claude Sonnet 5

#165

Earlier quoted context omitted.

Why do you think they are bragging? Anthropic has long been the company to give us by far the most in-depth information about their models, both positive and negative. I read this as them just stating a fact about this model that users would want to know.

The preceding sentence is >Our safety assessments found that Sonnet 5 shows an overall lower rate of undesirable behaviors than Sonnet 4.6, and is generally safer to use in agentic contexts. which is obviously painting that as a good thing. So reading the next sentence as "in other good news" is reasonable.

While I'm still not sure I would characterize that as bragging, you're right that that is a fair interpretation. However, another Fair interpretation of that is something along the lines of "the downside or cost of this positive thing is this following negative thing."

Re: Claude Sonnet 5

#167

Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models. I have been using Sonnet 4.6 more than Opus, because I'm mostly doing agent-assisted development and not fully agent-driven development. This announcement does not make me positive, I have fou…

From my own experience, GLM-5.2 generally cost more tokens and much more slow.

I use GLM 5.2 Fast from Fireworks and its very fast. Where are you using it from?

Re: Claude Sonnet 5

#168
Sonnet 5 is not currently available in the EU region on Bedrock, whereas previous models were and still are. I wonder if this is only due to early stages of the rollout or if this is due to recent US restrictions.

Unfortunately that means I won't be using it at work for now.

Re: Claude Sonnet 5

#170
> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models.

It seems being incompetent is a feature now...

Post reply on HN