Live data from Hacker News

Claude Sonnet 5

anthropic.com

181–190 of 822 posts

Re: Claude Sonnet 5

#181
The use of the "cheaper models" in big AI companies are next to useless as they don't even score as well as the open/super cheap Chinese models. Only the frontier big models like Fable and Opus have value.

Re: Claude Sonnet 5

#182
post #14

I didn't think they'd actually release a model that was worse than the open-weight frontier and at a higher price-point. Wow.

Why did the other reply to this get flagged as dead? It was a comment about how someone would come out saying that Sonnet 5 would be better on the pelican test and therefore it has to be good. But I guess HN loves pelican SVGs so much that you're not allowed to criticize it.

Re: Claude Sonnet 5

#184
The jump in reasoning quality is noticeable. What's interesting is how it handles ambiguous instructions now — it seems to ask fewer clarifying questions and just makes a reasonable judgment call. That's a double-edged sword depending on your use case.

Re: Claude Sonnet 5

#186

Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models. I have been using Sonnet 4.6 more than Opus, because I'm mostly doing agent-assisted development and not fully agent-driven development. This announcement does not make me positive, I have fou…

> I have been moving more and more to K2.7 Code and GLM-5.2 the last few weeks. They are often good enough for assistance, very fast, and cheap.

I've moved completely to local models that I run with my M1 Mac Studio (64gb ram) some time ago. But for the rare times when I feel the local, quantized Qwen3.6 isn't enough, I just connect to Openrouter and use something like Kimi, GLM or Deepseek for a fraction of the price of Anthropic et al.

Re: Claude Sonnet 5

#188

Earlier quoted context omitted.

Yeah, there's a real opportunity for one of these companies to invest time in a model that's tuned for, to use your term, agent-assisted developement. Trouble is, everyone inside their buildings seems to believe that no one will be working like that in a year or two.

There’s no way to justify their valuations if they get downgraded to a pair programming tool. They need fully agentic stuff to work and replace human engineers to even come close. Offhand, I’m not even certain whether a model like that could justify the constant retraining we’re doing on the agentic models. It doesn’t make a lot of sense to spend millions or billions on training to reduce hallucinations by 0.3% if yo…

My two cents is that the way to square this circle is that the valuations should be lower and they should be spending a lot less on constant retraining.

Unfortunately (from my perspective) it seems like the US companies are increasingly stuck in their current model. I think it's a competitive disadvantage.

But obviously most of the real insiders seem to disagree with me, so I'm probably wrong :)

Re: Claude Sonnet 5

#189

interesting how much worse the sentiment around Anthropic is getting

Seems like a combination of multiple factors:

"They took my shit away!" -- 3-day Fable 5 addicts (me)

"How dare they tell Trump no?" -- US nationalist / "my country right or wrong" types

"Great to see a closed source company fail!" -- open source boosters

"Great to see an American company fail!" -- anti-US, and/or pro-China folks

"Great to see a successful company fail!" -- anti-capitalists and/or sour-grapes crab bucket types

"Serves you right for ripping off creators!" -- copyright warriors

"They keep silently nerfing the models!" -- secret downgrade conspiracy theorists

"Quit killing the planet!" -- anti-datacenter advocates

Re: Claude Sonnet 5

#190

Earlier quoted context omitted.

Why do you think they are bragging? Anthropic has long been the company to give us by far the most in-depth information about their models, both positive and negative. I read this as them just stating a fact about this model that users would want to know.

Anthropomorphic, most in-depth? That's laughable given how closed down they've been over the years. If you want in-depth, DeepSeek actually still publishes papers of their methods for anyone to implement leading to being by far the most cost efficient model provider for the performance.

I was talking about reporting on testing and capabilities. Yes, open models provide a greater amount of information about the development of the model and how to run it yourself, but I am quite confident that literally no AI company, open or closed, conducts and reports so thoroughly on testing about the capabilities of their models.
Post reply on HN