Live data from Hacker News

Claude Sonnet 5

anthropic.com

711–720 of 822 posts

Re: Claude Sonnet 5

#711

Earlier quoted context omitted.

They have updated it

Did Anthropic have Opus 4.8 and Sonnet 5 switched in the Agentic Search chart at first?

No, and the original had everything more expensive. There's a comparison here:

https://www.reddit.com/r/ClaudeAI/comments/1ukgqwr/looks_lik...

The explanation Anthropic gave for the update doesn't address how the x-axis needed to range up to $50 previously and only $10 now. In any case the pass rates are also lower.

Probably the difference between whatever it is people notice when they say models become "nerfed".

Re: Claude Sonnet 5

#712

I'm struggling to understand why I'd ever use this instead of just using a lower effort level for opus given on many of the benchmarks listed the cost per task rises above opus at anything higher than medium effort. Only thing I can think of is for when someone is out of opus credits. Of course there are API billing use cases but I'd probably still just use opus on low.

I concur. I already use Opus 4.8 for almost all my tasks and this gives me almost no reason to try Sonnet 5.

Re: Claude Sonnet 5

#713
post #399

Claude Sonnet 5 itself described its pelican as looking like a goose: > Illustration of a white goose riding a bicycle, with one wing extended forward to grip the handlebar, set against a plain white background with a brown ground line. https://simonwillison.net/2026/Jun/30/claude-sonnet-5/

Mine is to ask it to write a parallel parking simulator and animation. The math there is surprisingly complex including differential equations. Fable 5 can almost one shot it with all tunable params.

Re: Claude Sonnet 5

#714
post #692

Earlier quoted context omitted.

More and more I find myself trying to stop Opus from doing something stupid, and at every turn I need to tell it to stop overcomplicating things. I think the models are being optimized for wealth extraction from users and companies, instead of solving problems. I don't know why Opus would try to create an entire library when I told it specifically to do something simple that would take 2-3 lines of Python.

> I think the models are being optimized for wealth extraction from users and companies, instead of solving problems. I don't think so. Expect that in a market with high vendor lock-in but that's not the case here. The market is extremely competitive and switching cost are near zero. Anthropic can't afford to pull shit like this and sacrifice quality.

The disconnect between the reality of and the consumer sentiment of this particular realm of products seems to be one of the most dramatic and widespread I’ve personally ever seen.

Re: Claude Sonnet 5

#715
post #301

Earlier quoted context omitted.

There is no nefariousness in owning all the means of production, it's the endgame of maximizing profit. However the result is exactly the same, concentration of power.

This is such a defeatist and low agency take. "means of production" are not a limited resource like gold that you have to extract from natural sources or divvy up. They are fundamentally skill and knowledge that anyone can attain and put to use, maybe not on the same scale as a well funded business but even those businesses had to start somewhere in order to grow to the size they are now. So rather than casting asper…

AI companies are trying to mechanize skill and knowledge and to own the infrastructure around it. If they succeed, your suggestion does not work. Even if they can't succeed, they will try because that's the most obvious path to maximizing profit for them.

Also about "creating" means of production, these companies actively try to sabotage this as another profit maximization strategy. They buy all the ram, so others cannot compete. They buy startups who succeed, so they stop competing.

It's not aspersions, it's just describing the phenomenon.

Even if I take your suggestion to heart, once my company would be big enough, if I wanted to optimize for profit, I would have to do the same as these companies.

The end result is concentrated power.

Re: Claude Sonnet 5

#716
post #704

Edit June 30, 2026: In the original version of this post, we included a cost-performance chart for the BrowseComp evaluation that was based on data from a simpler methodology that did not reflect the standard methodology we use for agentic search evaluations. This had the result of underestimating Sonnet 5's performance on the evaluation. They changed the Sonnet 5 'Agentic search' benchmark graph overnight

[deleted]

Re: Claude Sonnet 5

#717

I'm struggling to understand why I'd ever use this instead of just using a lower effort level for opus given on many of the benchmarks listed the cost per task rises above opus at anything higher than medium effort. Only thing I can think of is for when someone is out of opus credits. Of course there are API billing use cases but I'd probably still just use opus on low.

Is there a router or wrapper that provides a real-time cost estimation for alternative settings? Obviously, you can't predict exact output tokens without running the inference, but a tool that calculates the exact input cost across models and applies a historical average for the output tokens could be useful. Like, you run a task on Sonnet, and it estimates: "Based on your input tokens and a 1:1 output ratio, this would have cost $X on Opus at a low effort level."

Re: Claude Sonnet 5

#719
post #480
post #399

Claude Sonnet 5 itself described its pelican as looking like a goose: > Illustration of a white goose riding a bicycle, with one wing extended forward to grip the handlebar, set against a plain white background with a brown ground line. https://simonwillison.net/2026/Jun/30/claude-sonnet-5/

That's possibly the worst pelican I saw from all recent LLMs. Meanwhile GLM 5.2 drew a cool self-contained fully animated SVG pelican: https://simonwillison.net/2026/Jun/17/glm-52

Just need the legs to interact with the pedals now :D

Re: Claude Sonnet 5

#720

Seems to be another great incremental update to the workhorse, nice! I've been using Sonnet instead of Opus for almost all coding tasks for a while now. A little elbow grease to break down tasks and you can spend a lot less money for just about the same output quality.

It's a 30% price increase once the discount rate vanishes.
Post reply on HN