Live data from Hacker News

Claude Sonnet 5

anthropic.com

11–20 of 822 posts

Re: Claude Sonnet 5

#11
What I starting to hate is that each model's effort level can mean completely different power.

Today sonnet 5's med level effort is equivalent to sonnet 4.6 low level effort :/

Re: Claude Sonnet 5

#12
Seems to be another great incremental update to the workhorse, nice!

I've been using Sonnet instead of Opus for almost all coding tasks for a while now. A little elbow grease to break down tasks and you can spend a lot less money for just about the same output quality.

Re: Claude Sonnet 5

#13

Interesting that tasks on extra high cost almost the same as Opus 4.8 with a slightly worse performance

LRMs are plateauing for sure, not that there won't be gains to be had in the future, but it's not like the era of rapid progress that was the past year any more.

Re: Claude Sonnet 5

#15
post #8

That’s nice, but we want Fable

The reality is that Fable will eventually be obsolete and Sonnet / Opus will surpass it. Fable did cost 2x as much as Opus, so I assume it involves a much higher cost for what it did, but I wouldn't be surprised if Fable will be obsoleted by Opus or even Sonnet sooner or later at less cost.

Re: Claude Sonnet 5

#18
post #10

Opus 4.8 beats Sonnet 5 on the pareto frontier in several of their graphs (Agentic Search, Agentic Computer Use). In other words, for certain tasks, Opus 4.8 is cheaper than Sonnet 5, and does better than Sonnet 5. I've noticed this pattern on a lot of benchmarks. You can try to emulate a bigger model by ramping up the test time compute (max reasoning, more turns, model fusion etc.), but you can't reach the same qual…

And Claude Code penalizes you for using Sonnet on the subscription plan, so there's little reason to use it.

Re: Claude Sonnet 5

#19
Wow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters.

From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable than Opus 4.8 and Mythos 5

As with the other evaluations in this section, these results were achieved with all safeguards turned off. When run with our default mitigations, Sonnet 5 scored a 0 on CyberGym"

Re: Claude Sonnet 5

#20
post #11

What I starting to hate is that each model's effort level can mean completely different power. Today sonnet 5's med level effort is equivalent to sonnet 4.6 low level effort :/

[deleted]
Post reply on HN