Live data from Hacker News

Claude Sonnet 5

anthropic.com

71–80 of 822 posts

Re: Claude Sonnet 5

#71

Seems like the way to go for any smaller models is to only use the low reasoning levels, and for anything where you'd want it to reason harder, to just use a larger model. In effect, high reasoning only makes sense when you're using the frontier model and need extra performance (higher levels of reasoning are never pareto optimal unless you're at the largest model size).

Not to sound like an LLM, but that seems exactly right to me. Use it as a cheaper, high-functioning task subagent and lower reasoning for a master Opus session. As long as not every portion of your task requires maximum intelligence, you should come out ahead.

Re: Claude Sonnet 5

#72
post #18
post #10

Opus 4.8 beats Sonnet 5 on the pareto frontier in several of their graphs (Agentic Search, Agentic Computer Use). In other words, for certain tasks, Opus 4.8 is cheaper than Sonnet 5, and does better than Sonnet 5. I've noticed this pattern on a lot of benchmarks. You can try to emulate a bigger model by ramping up the test time compute (max reasoning, more turns, model fusion etc.), but you can't reach the same qual…

And Claude Code penalizes you for using Sonnet on the subscription plan, so there's little reason to use it.

This is what I realized, can you provide more detail on how you've observed this? The /usage screen does not make it clear.

Re: Claude Sonnet 5

#73
post #42

Earlier quoted context omitted.

They spent months hyping up Mythos and ended up with it banned. I’d assume they want to both differentiate their products and appeal to regulators here

They will release it eventually. Once they see the Chinese models are close to Mythos level they will release it before, so it will be "revolutionary".

It was already released. US government is the only reason it's not available to us mere mortals anymore

Re: Claude Sonnet 5

#75

> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. And Opus 4.8 is still cheaper for a higher pass rate (much less open weight models like GLM 5.2) so not sure why I'd use Sonnet except on the…

There's two classes of models now - the cybersecurity ones that none of us are getting, and the 'safe' models released for general consumption. This is letting us know which side of the divide it sits on.

Re: Claude Sonnet 5

#76
interesting footnotes: "Sonnet 5 is an upgrade to Sonnet 4.6, but it uses an updated tokenizer... can map to more tokens: roughly 1.0–1.35× depending on the content type." AKA expect higher costs on Sonnet 5 vs Sonnet 4.6 for the same tasks.

Re: Claude Sonnet 5

#77

> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. And Opus 4.8 is still cheaper for a higher pass rate (much less open weight models like GLM 5.2) so not sure why I'd use Sonnet except on the…

Why do you think they are bragging? Anthropic has long been the company to give us by far the most in-depth information about their models, both positive and negative. I read this as them just stating a fact about this model that users would want to know.

I'm absolutely certain that their marketing team has input on (if not owning) these announcements.

Re: Claude Sonnet 5

#79

Wow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters. From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable than Opus 4.8 and Mythos 5 As with the other evaluations in this section, these results were achieved with all safeguards turned off. When run with our default mitigations, Sonnet 5 scored a 0 on CyberG…

Not to single you out, parent commenter, but I really hope the quality of discourse on HN will move past these basic comparisons eventually. It seems like every thread on every model release has the exact same comments.

"Wow, X models is Y% better or worse than Claude Z model on T benchmark"

"That's irrelevant, they're just benchmaxing."

"Not useable for daily coding or agentic workloads, the vibes are totally wrong."

"It's almost as good, and costs a lot less, so I will absolutely use it."

"I cannot imagine justifying using these, as the step change means open models lower costs do not make up for the productivity loss"

I'm an unhappy Anthropic customer and really rooting for open models and non-gatekept intelligence, but how do we move on from this now meme-like model release discourse rigamarole. I do not know what that would be. I don't design LLMs nor benchmarks, and I genuinely appreciate that people do their best to provide information, even if non-perfect here. I'm sure most of you who actively read these comment pages on announcements must feel similarly, though, right?

Post reply on HN