Live data from Hacker News

Claude Sonnet 5

anthropic.com

51–60 of 822 posts

Re: Claude Sonnet 5

#51
post #32

> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. And Opus 4.8 is still cheaper for a higher pass rate (much less open weight models like GLM 5.2) so not sure why I'd use Sonnet except on the…

To avoid Lutnick getting on their case again.

He has the opportunity to do the funniest thing ever

Re: Claude Sonnet 5

#52
post #42

> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. And Opus 4.8 is still cheaper for a higher pass rate (much less open weight models like GLM 5.2) so not sure why I'd use Sonnet except on the…

They spent months hyping up Mythos and ended up with it banned. I’d assume they want to both differentiate their products and appeal to regulators here

They will release it eventually. Once they see the Chinese models are close to Mythos level they will release it before, so it will be "revolutionary".

Re: Claude Sonnet 5

#54

> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. And Opus 4.8 is still cheaper for a higher pass rate (much less open weight models like GLM 5.2) so not sure why I'd use Sonnet except on the…

so it doesn't get blocked. last time they said a model was great at cyber it didnt turn out well

Re: Claude Sonnet 5

#55
post #41
post #9

Earlier quoted context omitted.

This is on the browsercomp graph, right? In that, it seems sonnet 5 on high costs more than opus 4.8 at a lower pass rate. Am I reading this correctly? Edit: It looks like the key value proposition of the updated model is that it is much better than Sonnet 4.6. Wheras, Sonnet 5 delivers great value (by browsercomp benchmarks and compared to opus) when running in low and medium. So: Sonnet 4.6 should ~never have been…

I agree with this assessment, IMO my takeaway from this is "Generally run Sonnet on low, otherwise use Opus". It's kind of like an "extra low" setting of Opus. (depends on the application for sure).

It would be good if Anthropic provided some kind of feedback or even toggle to auto-route requests for models being used at thinking levels that would be a better value using a different model.

Sort of like, getting an automatic upgrade at a car rental or hotel if there is availability.

Re: Claude Sonnet 5

#56

The cost per task chart is telling me that I should _never_ use Sonnet 5 above medium effort level - Opus always performs better for a given cost. So I guess the takeaway is that if Sonnet 5 medium isn't good enough for you, switch models, not effort levels.

Yeah, I was looking at the same chart and was very surprised at where the curve is relative to opus... Feels like sonnet 5 is "what if opus had an extra-low effort level"?

Re: Claude Sonnet 5

#58
Anybody notice that they did not include Sonnet 5 Max in the "Agentic Search results", when comparing to Opus 4.8 ...

Based upon the "Agentic Computer usage", Sonnet 5 Max was going to be off "Agentic Search results" chart. lol ...

In short, Sonnet 5 Low/Medium is more cost efficient, if its a task below Opus 4.8 Medium. For the rest its expensive and your better off using Opus 4.8.

Why even release this model?

Re: Claude Sonnet 5

#59
post #47

Wow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters. From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable than Opus 4.8 and Mythos 5 As with the other evaluations in this section, these results were achieved with all safeguards turned off. When run with our default mitigations, Sonnet 5 scored a 0 on CyberG…

Finally, a viable business strategy - sell security-oblivious code monkeys for cheap, then charge premium rates for agents capable of cleaning up the mess.

I think instead they should sell super hackers and get their product banned instantly and go bankrupt
Post reply on HN