Today sonnet 5's med level effort is equivalent to sonnet 4.6 low level effort :/
Claude Sonnet 5
11–20 of 822 posts
Re: Claude Sonnet 5
#12I've been using Sonnet instead of Opus for almost all coding tasks for a while now. A little elbow grease to break down tasks and you can spend a lot less money for just about the same output quality.
Re: Claude Sonnet 5
#13Interesting that tasks on extra high cost almost the same as Opus 4.8 with a slightly worse performance
Re: Claude Sonnet 5
#14Re: Claude Sonnet 5
#15That’s nice, but we want Fable
Re: Claude Sonnet 5
#16Re: Claude Sonnet 5
#17I didn't think they'd actually release a model that was worse than the open-weight frontier and at a higher price-point. Wow.
Re: Claude Sonnet 5
#18Opus 4.8 beats Sonnet 5 on the pareto frontier in several of their graphs (Agentic Search, Agentic Computer Use). In other words, for certain tasks, Opus 4.8 is cheaper than Sonnet 5, and does better than Sonnet 5. I've noticed this pattern on a lot of benchmarks. You can try to emulate a bigger model by ramping up the test time compute (max reasoning, more turns, model fusion etc.), but you can't reach the same qual…
Re: Claude Sonnet 5
#19From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable than Opus 4.8 and Mythos 5
As with the other evaluations in this section, these results were achieved with all safeguards turned off. When run with our default mitigations, Sonnet 5 scored a 0 on CyberGym"
Re: Claude Sonnet 5
#20What I starting to hate is that each model's effort level can mean completely different power. Today sonnet 5's med level effort is equivalent to sonnet 4.6 low level effort :/