I also like that the difference between low, medium, high, xhigh seems more spread, which is actually a good thing for people trying to tune applications. Running Sonnet 5 on low with the launch pricing makes this potentially a better fit than Haiku or open source models for some tasks. I don't think it will make sense at full price.
Claude Sonnet 5
21–30 of 822 posts
Re: Claude Sonnet 5
#22Re: Claude Sonnet 5
#23Re: Claude Sonnet 5
#24Re: Claude Sonnet 5
#25Interesting that tasks on extra high cost almost the same as Opus 4.8 with a slightly worse performance
LRMs are plateauing for sure, not that there won't be gains to be had in the future, but it's not like the era of rapid progress that was the past year any more.
Re: Claude Sonnet 5
#26Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability.
And Opus 4.8 is still cheaper for a higher pass rate (much less open weight models like GLM 5.2) so not sure why I'd use Sonnet except on the low effort level for I suppose trivial tasks where I want it to work only 50% of the time judging by the graph. The pricing doesn't really make any sense.
Re: Claude Sonnet 5
#27Re: Claude Sonnet 5
#28In effect, high reasoning only makes sense when you're using the frontier model and need extra performance (higher levels of reasoning are never pareto optimal unless you're at the largest model size).