Live data from Hacker News

Claude Sonnet 5

anthropic.com

201–210 of 822 posts

Re: Claude Sonnet 5

#201

The cost per task chart is telling me that I should _never_ use Sonnet 5 above medium effort level - Opus always performs better for a given cost. So I guess the takeaway is that if Sonnet 5 medium isn't good enough for you, switch models, not effort levels.

Worth noting that the default chart there is for "agentic search performance", not coding. I didn't see an effort comparison for coding specifically.

Re: Claude Sonnet 5

#202
Important to note: "Sonnet 5 is an upgrade to Sonnet 4.6, but it uses an updated tokenizer that changes how the model processes text to improve performance (this is similar to the tokenizer change we introduced with Claude Opus 4.7). The tradeoff is that the same input can map to more tokens: roughly 1.0–1.35× depending on the content type. The introductory pricing is set so that the transition to Sonnet 5 is roughly cost-neutral."

Re: Claude Sonnet 5

#203

I'd love if they would include speed (though I know there are difficulties involved). At this point the quality of Opus 4.8 is no longer my limiting factor, it's the speed, so a faster model would be great.

Have you tried Opus on fast mode?

Re: Claude Sonnet 5

#204

interesting how much worse the sentiment around Anthropic is getting

Seems like a combination of multiple factors: "They took my shit away!" -- 3-day Fable 5 addicts (me) "How dare they tell Trump no?" -- US nationalist / "my country right or wrong" types "Great to see a closed source company fail!" -- open source boosters "Great to see an American company fail!" -- anti-US, and/or pro-China folks "Great to see a successful company fail!" -- anti-capitalists and/or sour-grapes crab bu…

It seems to be more them losing goodwill combined with their marketing.

I don't agree with your framing that all negativity is from crazies

Re: Claude Sonnet 5

#205

Earlier quoted context omitted.

What? If you're comparing their models in the same size class, Sonnet 5 is Pareto-optimal over Sonnet 4.6.

I think they mean per dollar in the perf/$charts, not per marketing class. I.e. the new model is a complete Pareto failure in said perf/$ charts with the sole exception of Sonnet 5 low, which is dumb enough to not have comparison at all. Opus 4.8 delivers a better outcome per dollar, regardless what the underlying size of the models is. I'd generously assume this is something about the specific category of agentic ta…

For agentic computer use Sonnet 5 low performs better than Sonnet 4.6 medium at just under half the cost, and better than Opus 4.8 low at 25% off. Their success rates are not that far off.

Agentic search is a different story, but even there it still dominates 4.6 (as in, for everything Sonnet 4.6 can do, Sonnet 5 can do it as well or better at the same or lower cost).

Yes, Opus 4.8 dominates Sonnet 5 over its entire range in both categories, but Opus's lower range is limited and there is a valid regime on the lower end where Sonnet 5 use makes economic sense. This is not the case for Sonnet 4.6 where Opus 4.8 dominates it completely on both charts.

Edit -- reading your response closer I think we're saying the same things, maybe just disagreeing on whether that lower end is valuable or not.

Re: Claude Sonnet 5

#206

Not looking great for an upcoming IPO

You’re right, it’s looking stellar. Well beyond great. Real, and unprecedented, revenue growth will do that for a company.

"Real and unprecedented revenue growth"

Bro that is financial engineering, not real revenue growth. They engineered the switch to usage based pricing and a price hike timed the quarter before they wanted to go public, long enough to juice their numbers but not long enough for them not to be able to manage backlash and have to walk things back. Then they tried to extrapolate that manufactured bump to make it look like they have record shattering revenue growth.

Re: Claude Sonnet 5

#208

The cost per task chart is telling me that I should _never_ use Sonnet 5 above medium effort level - Opus always performs better for a given cost. So I guess the takeaway is that if Sonnet 5 medium isn't good enough for you, switch models, not effort levels.

What is a "task" in real-world terms? If it will be $15/million output tokens, and high/xhigh is somewhere in the $7.50/task range. Does that mean a single task is using 500k tokens. That seems like it would start to add up fast.

Re: Claude Sonnet 5

#209

Earlier quoted context omitted.

Seems like a combination of multiple factors: "They took my shit away!" -- 3-day Fable 5 addicts (me) "How dare they tell Trump no?" -- US nationalist / "my country right or wrong" types "Great to see a closed source company fail!" -- open source boosters "Great to see an American company fail!" -- anti-US, and/or pro-China folks "Great to see a successful company fail!" -- anti-capitalists and/or sour-grapes crab bu…

It seems to be more them losing goodwill combined with their marketing. I don't agree with your framing that all negativity is from crazies

I don't think all the negativity is from crazies, but big chunks of it are certainly motivated. I certainly left out numerous other categories.

Re: Claude Sonnet 5

#210

Earlier quoted context omitted.

Okay I don’t care about “eventually”, I want Fable now.

Have you considered getting better at coding so you can build stuff yourself instead of waiting for models you might not be able to get access to anymore?

This is like telling someone who wants a motorcycle that they should get better at running instead.
Post reply on HN