Live data from Hacker News

Claude Sonnet 5

anthropic.com

421–430 of 822 posts

Re: Claude Sonnet 5

#421
post #202

Important to note: "Sonnet 5 is an upgrade to Sonnet 4.6, but it uses an updated tokenizer that changes how the model processes text to improve performance (this is similar to the tokenizer change we introduced with Claude Opus 4.7). The tradeoff is that the same input can map to more tokens: roughly 1.0–1.35× depending on the content type. The introductory pricing is set so that the transition to Sonnet 5 is roughly…

So the post-introductory price is set such that Sonnet 5 will cost 100%-135% as much?

Correct. Albeit the nuance here is that a more capable model might solve problems more efficiently and faster, possibly saving you tokens.

As with any new model, you won't know the real impact until you start using it for your workload.

Re: Claude Sonnet 5

#422
post #395

Earlier quoted context omitted.

As always, note: faster than GLM-5.2 doesn't mean too much, as GLM-5.2 is served by different providers, so the inference speed can vary drastically between providers or over time.

What’s everyone favorite GLM provider? z.ai doesnt always have the most reliable AI but I don’t mind the party seeing my trade secrets and thoughts compared to an American corporation + the party seeing my trade secrets and thoughts. So thats not a functional difference to me, and the Chinese one won’t reply to subpoenas so thats a value add tbh So I’ll consider all, fastest tokens/sec wins

Fireworks.ai is solid. And if you care more about speed than cost they have a "fast" variant that I think just throws more hardware at the model for about 2x the cost.

Re: Claude Sonnet 5

#423
post #386

Earlier quoted context omitted.

I really don't get what you're proposing. The cost ranges do not overlap at the low end. You can't (by definition!) interpolate outside of the range. If you mean extrapolate, at that point you're just making up data. The available effort levels are discrete and covered totally by the benchmarks. You can draw on the monitor with a sharpie to show a "ultra-low" effort level for Opus that scores better than Sonnet "low"…

That's why I said "over the shared frontier" in my first post and more precisely in my second post I said "over the overlapping x values for which both are defined." It was a claim that applies to a range of x-values where both curves are defined. Of course if you go beyond those x-values where only one of the two are defined, then trivially the one that is defined constitutes the Pareto frontier in that region. Whic…

The post I was replying to said "performs strictly better at the same cost per task". That claim was obviously not true, there are costs where Opus cannot do the task and Sonnet can, so Opus can't be performing strictly better that the same cost. It seems that you agree that it is not true.

You could make it true by artificially dropping some of the data points, but, like, why?

(Again, this is moot given the updated graph.)

> Of course if you go beyond those x-values where only one of the two are defined, then trivially the one that is defined constitutes the Pareto frontier in that region.

Not so! It's only sound to do that at the low end of the cost axis (x) or the high end of the performance axis (y). You can't do it at the low end of the performance axis or the high end of the cost axis.

Re: Claude Sonnet 5

#424

Earlier quoted context omitted.

There’s no way to justify their valuations if they get downgraded to a pair programming tool. They need fully agentic stuff to work and replace human engineers to even come close. Offhand, I’m not even certain whether a model like that could justify the constant retraining we’re doing on the agentic models. It doesn’t make a lot of sense to spend millions or billions on training to reduce hallucinations by 0.3% if yo…

That's a really good point. I think if there wasn't the insane amount of money involved and these were treated as tools instead, they would probably be MORE productive. I think a person working hand in hand with an AI instead of delegating is the sweet spot of making things fast while also not losing understanding or control of the system. You are absolutely right that these companies can't justify their valuations i…

> I'm thinking the future of this tech will likely be better tooling with better IDE integrations rather than "Claude plz make me a SaaS kthx"

I think this sort of thinking is a trap, because it presumes that all software has the same constraints.

There's a spectrum of requirements between "chuck this over the wall at Claude, it only has to work once" and "this is a literal rocket ship, formally verify the whole thing".

I've made some things with Claude I don't understand and don't control. It's fine, they're still useful to me. Things for the house that I wasn't going to build manually, some dashboarding stuff and scripts for work, stuff that can crash and burn and I'll be fine.

They won't justify trillions in investment, but they are useful.

Equally, I do agree with you on some things. Sometimes I hand-hold the LLM or forgo it entirely because I want to be 100% sure I know how something works, and can justify a decision if it causes a production outage.

I think the future is probably multiple different tools with different goals. Better IDE integration for some uses, an entirely separate "LLM herd controller" kind of thing for when you're okay with vibe-coding, and the most interesting is something in the middle where you're more in the loop than pure vibe-coding, but don't see the full context like in an IDE. Something where it surfaces changes to key components, but hides things like test changes.

Re: Claude Sonnet 5

#425
post #65
post #45

Wonder if the whole cyber paranoia leads to their models ultimately generating less secure code. After all, if it has the ability to generate safe code, it would imply that it knows something about cybersecurity, which could surely be used to hack all the banks in the world.

Trying to censor nudity in image generation models caused all kinds of problems with anatomy in image models. I’m sure these models will have similar issues with security.

Censorship on image generation models works on another level. The models can generate NSFW, but there are extra computer vision models checking if the images can be shown to the users. It's especially obvious for Grok and ChatGPT.

Re: Claude Sonnet 5

#426

The cost per task chart is telling me that I should _never_ use Sonnet 5 above medium effort level - Opus always performs better for a given cost. So I guess the takeaway is that if Sonnet 5 medium isn't good enough for you, switch models, not effort levels.

While I appreciate, they publish this information, it's increasingly hard to keep track of it all. I've lost the mental model of how different models at different effort levels perform and what tasks they are good at. In practice, I tend to just use the default on Claude Code that works well enough. But I wonder to what degree other users really play around with these settings to optimize for their project.

Just because it’s hard to keep track of doesn’t mean it’s not relevant.

Playing around with learning the differences is incredibly helpful to schedule on ones calendar weekly for an hour or two, while saving links throughout the week to try out.

Re: Claude Sonnet 5

#427

Earlier quoted context omitted.

I don't think all the negativity is from crazies, but big chunks of it are certainly motivated. I certainly left out numerous other categories.

The amount of anti-Anthropic and anti-Dario posts i've seen on reddit threads has gotten a bit ridiculous. It feels like your analysis is mostly spot on, it's the confluence of several motivated parties pouring effort into social media. Many of the posters are pro-foreign models/pro-open source, and most can't distinguish the difference between "open source" and open weight models like Qwen, Minimax, or GLM. Reminds…

Fable as released was censored to the point of being useless for many tasks. Now surprise surprise it's not even available unless you're pre-approved.

Qwen is also censored - although since it's open weight, there are completely uncensored versions available.

The owners of Qwen can't jack up the prices to something I'm unable to pay. They can't take it away.

The owners of Qwen can't log and train on my data.

Open weight models share far more in common with free speech than free beer.

If big daddy Dario and his company are getting pushback it's not being of some motivated group trying to take them down. They brought it on themselves.

Re: Claude Sonnet 5

#428

The cost per task chart is telling me that I should _never_ use Sonnet 5 above medium effort level - Opus always performs better for a given cost. So I guess the takeaway is that if Sonnet 5 medium isn't good enough for you, switch models, not effort levels.

it might be worth it if speed is an issue

Re: Claude Sonnet 5

#429

Earlier quoted context omitted.

I'd love to meet the devs who can spin up full feature web apps in under 15 minutes with all the bells and whistles I've gotten Claude to spin up and code. I don't think the AI haters understand the level of time cutting that you can achieve with a very simple and reasonably crafted prompt. I'm talking back-end, with database models, classes, queries, accompanying front-end layouts, with real dynamic data, running. S…

And the trade off for that productivity is relying on a completely untrustworthy company/product that gets more expensive and uncertain by the week while your skills erode.

Companies don't care about your skillz, they care about velocity and costs. If AI helps increase velocity and decrease cost by lowering total headcount, then its a massive win. That factors in AI "unpredictability".
Post reply on HN