Live data from Hacker News

Claude Sonnet 5

anthropic.com

41–50 of 822 posts

Re: Claude Sonnet 5

#41
post #9

Interesting that tasks on extra high cost almost the same as Opus 4.8 with a slightly worse performance

This is on the browsercomp graph, right? In that, it seems sonnet 5 on high costs more than opus 4.8 at a lower pass rate. Am I reading this correctly? Edit: It looks like the key value proposition of the updated model is that it is much better than Sonnet 4.6. Wheras, Sonnet 5 delivers great value (by browsercomp benchmarks and compared to opus) when running in low and medium. So: Sonnet 4.6 should ~never have been…

I agree with this assessment, IMO my takeaway from this is "Generally run Sonnet on low, otherwise use Opus". It's kind of like an "extra low" setting of Opus. (depends on the application for sure).

Re: Claude Sonnet 5

#42

> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. And Opus 4.8 is still cheaper for a higher pass rate (much less open weight models like GLM 5.2) so not sure why I'd use Sonnet except on the…

They spent months hyping up Mythos and ended up with it banned. I’d assume they want to both differentiate their products and appeal to regulators here

Re: Claude Sonnet 5

#45
Wonder if the whole cyber paranoia leads to their models ultimately generating less secure code. After all, if it has the ability to generate safe code, it would imply that it knows something about cybersecurity, which could surely be used to hack all the banks in the world.

Re: Claude Sonnet 5

#47

Wow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters. From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable than Opus 4.8 and Mythos 5 As with the other evaluations in this section, these results were achieved with all safeguards turned off. When run with our default mitigations, Sonnet 5 scored a 0 on CyberG…

Finally, a viable business strategy - sell security-oblivious code monkeys for cheap, then charge premium rates for agents capable of cleaning up the mess.

Re: Claude Sonnet 5

#49

> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. And Opus 4.8 is still cheaper for a higher pass rate (much less open weight models like GLM 5.2) so not sure why I'd use Sonnet except on the…

[deleted]

Re: Claude Sonnet 5

#50

Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models. I have been using Sonnet 4.6 more than Opus, because I'm mostly doing agent-assisted development and not fully agent-driven development. This announcement does not make me positive, I have fou…

From my own experience, GLM-5.2 generally cost more tokens and much more slow.
Post reply on HN