Live data from Hacker News

Claude Sonnet 5

anthropic.com

1–10 of 822 posts

Re: Claude Sonnet 5

#9

Interesting that tasks on extra high cost almost the same as Opus 4.8 with a slightly worse performance

This is on the browsercomp graph, right?

In that, it seems sonnet 5 on high costs more than opus 4.8 at a lower pass rate. Am I reading this correctly?

Edit: It looks like the key value proposition of the updated model is that it is much better than Sonnet 4.6.

Wheras, Sonnet 5 delivers great value (by browsercomp benchmarks and compared to opus) when running in low and medium.

So: Sonnet 4.6 should ~never have been run for low, medium or high when Opus 4.8 has been available. Whoops, I think I have some skills that delegate easy stuff to Sonnet.

---

I remember Anthropic pivoting everyone's default model to Opus but had not seen it put so starkly before.

I am a bit confused on the subscription `/usage` screen. It splits out sonnet usage, and I'd presumed that would have contributed to a lower use of subscription Quota.

But if this is correct, Sonnet usage was basically like smoking unfiltered cigarettes.

Re: Claude Sonnet 5

#10
Opus 4.8 beats Sonnet 5 on the pareto frontier in several of their graphs (Agentic Search, Agentic Computer Use).

In other words, for certain tasks, Opus 4.8 is cheaper than Sonnet 5, and does better than Sonnet 5.

I've noticed this pattern on a lot of benchmarks. You can try to emulate a bigger model by ramping up the test time compute (max reasoning, more turns, model fusion etc.), but you can't reach the same quality level, and you often exceed the cost you would have paid by just using a bigger model.

tldr: if you're doing something hard, just use a bigger model.

Post reply on HN