> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. And Opus 4.8 is still cheaper for a higher pass rate (much less open weight models like GLM 5.2) so not sure why I'd use Sonnet except on the…
To avoid Lutnick getting on their case again.
Claude Sonnet 5
51–60 of 822 posts
Re: Claude Sonnet 5
#52> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. And Opus 4.8 is still cheaper for a higher pass rate (much less open weight models like GLM 5.2) so not sure why I'd use Sonnet except on the…
They spent months hyping up Mythos and ended up with it banned. I’d assume they want to both differentiate their products and appeal to regulators here
Re: Claude Sonnet 5
#53there was a vibecoded prediction market–style page that was put up yesterday (?) that got the date exactly right i think
link?
Re: Claude Sonnet 5
#54> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. And Opus 4.8 is still cheaper for a higher pass rate (much less open weight models like GLM 5.2) so not sure why I'd use Sonnet except on the…
Re: Claude Sonnet 5
#55Earlier quoted context omitted.
This is on the browsercomp graph, right? In that, it seems sonnet 5 on high costs more than opus 4.8 at a lower pass rate. Am I reading this correctly? Edit: It looks like the key value proposition of the updated model is that it is much better than Sonnet 4.6. Wheras, Sonnet 5 delivers great value (by browsercomp benchmarks and compared to opus) when running in low and medium. So: Sonnet 4.6 should ~never have been…
I agree with this assessment, IMO my takeaway from this is "Generally run Sonnet on low, otherwise use Opus". It's kind of like an "extra low" setting of Opus. (depends on the application for sure).
Sort of like, getting an automatic upgrade at a car rental or hotel if there is availability.
Re: Claude Sonnet 5
#56The cost per task chart is telling me that I should _never_ use Sonnet 5 above medium effort level - Opus always performs better for a given cost. So I guess the takeaway is that if Sonnet 5 medium isn't good enough for you, switch models, not effort levels.
Re: Claude Sonnet 5
#57Re: Claude Sonnet 5
#58Based upon the "Agentic Computer usage", Sonnet 5 Max was going to be off "Agentic Search results" chart. lol ...
In short, Sonnet 5 Low/Medium is more cost efficient, if its a task below Opus 4.8 Medium. For the rest its expensive and your better off using Opus 4.8.
Why even release this model?
Re: Claude Sonnet 5
#59Wow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters. From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable than Opus 4.8 and Mythos 5 As with the other evaluations in this section, these results were achieved with all safeguards turned off. When run with our default mitigations, Sonnet 5 scored a 0 on CyberG…
Finally, a viable business strategy - sell security-oblivious code monkeys for cheap, then charge premium rates for agents capable of cleaning up the mess.