Interesting that tasks on extra high cost almost the same as Opus 4.8 with a slightly worse performance
This is on the browsercomp graph, right? In that, it seems sonnet 5 on high costs more than opus 4.8 at a lower pass rate. Am I reading this correctly? Edit: It looks like the key value proposition of the updated model is that it is much better than Sonnet 4.6. Wheras, Sonnet 5 delivers great value (by browsercomp benchmarks and compared to opus) when running in low and medium. So: Sonnet 4.6 should ~never have been…
Claude Sonnet 5
41–50 of 822 posts
Re: Claude Sonnet 5
#42> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. And Opus 4.8 is still cheaper for a higher pass rate (much less open weight models like GLM 5.2) so not sure why I'd use Sonnet except on the…
Re: Claude Sonnet 5
#43Re: Claude Sonnet 5
#44When can we get a new Haiku? 4.5 came out nearly a year ago, and it's showing its age.
Re: Claude Sonnet 5
#45Re: Claude Sonnet 5
#46Okay.
Re: Claude Sonnet 5
#47Wow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters. From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable than Opus 4.8 and Mythos 5 As with the other evaluations in this section, these results were achieved with all safeguards turned off. When run with our default mitigations, Sonnet 5 scored a 0 on CyberG…
Re: Claude Sonnet 5
#48Earlier quoted context omitted.
[flagged]
"very aesthetically pleasing beak. good form. looks to be riding fast. please visit my website"
Re: Claude Sonnet 5
#49> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. And Opus 4.8 is still cheaper for a higher pass rate (much less open weight models like GLM 5.2) so not sure why I'd use Sonnet except on the…
Re: Claude Sonnet 5
#50Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models. I have been using Sonnet 4.6 more than Opus, because I'm mostly doing agent-assisted development and not fully agent-driven development. This announcement does not make me positive, I have fou…