Live data from Hacker News

Claude Sonnet 5

anthropic.com

241–250 of 822 posts

Re: Claude Sonnet 5

#243

Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models. I have been using Sonnet 4.6 more than Opus, because I'm mostly doing agent-assisted development and not fully agent-driven development. This announcement does not make me positive, I have fou…

No kidding. I expect to have models to use which are optimised for different use cases.

Sonnet as an autonomous agentic model is silly. We already have other models for that if you want something weaker and cheaper than Opus.

Re: Claude Sonnet 5

#244

Earlier quoted context omitted.

While I appreciate, they publish this information, it's increasingly hard to keep track of it all. I've lost the mental model of how different models at different effort levels perform and what tasks they are good at. In practice, I tend to just use the default on Claude Code that works well enough. But I wonder to what degree other users really play around with these settings to optimize for their project.

What I want is a harness that knows how to optimize this kind of thing for me.

You might want to check out Amp: https://ampcode.com/

Re: Claude Sonnet 5

#245

I'm struggling to understand why I'd ever use this instead of just using a lower effort level for opus given on many of the benchmarks listed the cost per task rises above opus at anything higher than medium effort. Only thing I can think of is for when someone is out of opus credits. Of course there are API billing use cases but I'd probably still just use opus on low.

More and more I find myself trying to stop Opus from doing something stupid, and at every turn I need to tell it to stop overcomplicating things.

I think the models are being optimized for wealth extraction from users and companies, instead of solving problems.

I don't know why Opus would try to create an entire library when I told it specifically to do something simple that would take 2-3 lines of Python.

Re: Claude Sonnet 5

#246

Earlier quoted context omitted.

There’s no way to justify their valuations if they get downgraded to a pair programming tool. They need fully agentic stuff to work and replace human engineers to even come close. Offhand, I’m not even certain whether a model like that could justify the constant retraining we’re doing on the agentic models. It doesn’t make a lot of sense to spend millions or billions on training to reduce hallucinations by 0.3% if yo…

My two cents is that the way to square this circle is that the valuations should be lower and they should be spending a lot less on constant retraining. Unfortunately (from my perspective) it seems like the US companies are increasingly stuck in their current model. I think it's a competitive disadvantage. But obviously most of the real insiders seem to disagree with me, so I'm probably wrong :)

The insiders disagree because they are benefiting greatly from the insane valuations, right?

Chinese models are quickly commodifying frontier inference, the US Gov is preventing domestic SOTA models access to the public and without those models why would consumers still spend $200/month to use the best models?

It’s such a mess and isn’t inspiring confidence as a non-investor.

Re: Claude Sonnet 5

#247

Interesting that tasks on extra high cost almost the same as Opus 4.8 with a slightly worse performance

LRMs are plateauing for sure, not that there won't be gains to be had in the future, but it's not like the era of rapid progress that was the past year any more.

I agree that the rapid improvement from like 2023-24 era is over (from a perspective of going from a 3/10 to a 7/10, you can’t then go to a 11/10). There was just so much more space to grow back then.

But isn’t Fable supposed to be another step change? I never used it, myself.

Tbh, at this point I think top tier models are smart “enough” (I’m sure this will look antiquated in a year), and the way to give me MORE noticeable improvement is to make them much faster rather than much smarter. Or even a way to automatically and accurately pick faster models when it makes sense. I know that IDE’s have Auto modes, but it’s not something that I trust right now to pick smart+fast instead of picking “maybe smart enough”+”cheaper for harness owner”

Re: Claude Sonnet 5

#248
post #182
post #14

I didn't think they'd actually release a model that was worse than the open-weight frontier and at a higher price-point. Wow.

Why did the other reply to this get flagged as dead? It was a comment about how someone would come out saying that Sonnet 5 would be better on the pelican test and therefore it has to be good. But I guess HN loves pelican SVGs so much that you're not allowed to criticize it.

If you look at the account history, it's pretty clearly an account-level thing, not a comment-level thing.

Re: Claude Sonnet 5

#250
post #44

Earlier quoted context omitted.

Look at Qwen for that level of intelligence.

needs to be on bedrock for me to use it at work

Gemma 4, Kimi K2.5, MiniMax M2.5, gpt-oss, GLM 5, Qwen3 Coder Next, DeepSeek V3.2, Devstral 2, are all available on AWS Bedrock and all are about Haiku level
Post reply on HN