Live data from Hacker News

Claude Sonnet 5

anthropic.com

701–710 of 822 posts

Re: Claude Sonnet 5

#701
post #227

Earlier quoted context omitted.

Older Opus models will likely get deprecated and then over time this is the cheapest model. That is how prices are currently increased.

Wat. Price/perf has been going down massively over the last few years.

Because they still haven't fully captured the market for Agentic Development.

Re: Claude Sonnet 5

#702

I don't know what I am doing right, or wrong, but I have access to claude and codex and I find myself giving the more serious work to codex recently. I tend to trust it more. I might try again Fable when it's back, but this Sonnet 5 didn't work well for my current projects.

Same (Opus 4.8 vs gpt 5.5)

I keep having to correct 4.8, but 5.5 more often than not is correcting me.

Opus writes a bit nicer though and it is easier to follow wat it is doing/saying. Not too different experience from talking to humans: 5.5 feels like a very smart 'nerd' that doesn't make a huge effort to communicate wel, while Opus is a bit less intelligent but that makes it's ideas easier to communicate

Re: Claude Sonnet 5

#703

Earlier quoted context omitted.

Are we reading the same chart? They have Sonnet You have to test each task obviously but it is not a bad model on its face.

They have updated it

Did Anthropic have Opus 4.8 and Sonnet 5 switched in the Agentic Search chart at first?

Re: Claude Sonnet 5

#704
Edit June 30, 2026: In the original version of this post, we included a cost-performance chart for the BrowseComp evaluation that was based on data from a simpler methodology that did not reflect the standard methodology we use for agentic search evaluations. This had the result of underestimating Sonnet 5's performance on the evaluation.

They changed the Sonnet 5 'Agentic search' benchmark graph overnight

Re: Claude Sonnet 5

#705
post #65

Earlier quoted context omitted.

Trying to censor nudity in image generation models caused all kinds of problems with anatomy in image models. I’m sure these models will have similar issues with security.

Censorship on image generation models works on another level. The models can generate NSFW, but there are extra computer vision models checking if the images can be shown to the users. It's especially obvious for Grok and ChatGPT.

That's only correct for specific models and not what parent was referring to.

Stable Diffusion 3, an open weights model, was laughed at at release for not being able to even generate a woman laying in grass. The community attributed this to the heavy dataset filtering. Since then other open weights releases have been made with no NSFW capabilities and the community claims they're not as good as anatomy as well.

You can google "stable diffusion 3 woman in grass" and press the images tab to see how the model failed spectacularly.

Re: Claude Sonnet 5

#706
post #666

Earlier quoted context omitted.

The skilled seniors better stop downplaying what actually led them to be skilled in the first place, and realize that the conditions to develop that skill has been gone and almost deemed unproductive in today's workplace. Not disagreeing that LLM's are a force multiplier, but I highly doubt whatever value will end up finding multiplying in the next generation of seniors, at this rate. It's surreal to me that I have t…

I have been thinking about it - paid apprenticeship is the answer - just as it is in other professions like medicine. Seniors should be paid to actively introduce juniors to the trade over couple of years. No more bootcamp entry. And it would be significant $ for senior to agree to expend his time and energy on software engineering apprentices. There would be also very limited number of places with good seniors. Exac…

> Seniors should be paid to actively introduce juniors

Instead lets train the contractors of an IT sourcing company and then we don't need you.

Re: Claude Sonnet 5

#707
post #132

Earlier quoted context omitted.

I have tried to rewrite an article with GLM-5.2 and with Sonnet 4.6. Completely different results as LLM is non-deterministic. But GLM-5.2 made a lot of subtle mistakes that needed to be corrected by hand. On the opposite, Sonnet found and corrected all mistakes in the second round. Similar situation was with planning and coding. GLM-5.2 seems to be good “on paper” but the real usage results was different. And I am n…

> Completely different results as LLM is non-deterministic. You'd need to produce this like 20 times by each model and then do 2x20x20 cross comparisons by both models and ultimately distill the 2x20x20 comparison results into two reports of how they differ. In this non deterministic computing future, everything else is voodoo, feelings and "vibes".

I would expect a model's result each time to be of a similar quality to the other times. There's something wrong if it does a way better or worse job, at the same problem, sometimes. It's possible, but I haven't heard anyone saying that they do.

Re: Claude Sonnet 5

#708
post #14

I didn't think they'd actually release a model that was worse than the open-weight frontier and at a higher price-point. Wow.

Today I tested GLM 5.2 by giving it an example stylesheet and told it to change the background color of a submit button.

It then hallucinated the submit button class...

Re: Claude Sonnet 5

#709

Earlier quoted context omitted.

> I don't think they're a net gain if you're a skilled senior I'm a skilled senior (I'm 54 and been coding since I was about 8; I've been 100% AI-generated code for at least 6 months now and have produced a combination of speed and quality that has astonished me; my velocity is apparent at https://github.com/pmarreck/ ) and this has been a massive net gain, so your claim is now officially in sheer defiance of reality…

I think I might have written a comment similar to yours maybe 6 months or a year ago. I'm not quite sure to respond to these sorts of replies. I have used LLMs/Claude Code quite extensively professionally and was a very early adopter, have built tooling around LLM/agentic development, and genuinely embraced it. They aren't useless, but the short term gains you think you're getting come at a very steep price that you…

> very steep price

I have yet to see it, but OK

Either measure it or it sounds like a conspiracy theory

Re: Claude Sonnet 5

#710

I'm struggling to understand why I'd ever use this instead of just using a lower effort level for opus given on many of the benchmarks listed the cost per task rises above opus at anything higher than medium effort. Only thing I can think of is for when someone is out of opus credits. Of course there are API billing use cases but I'd probably still just use opus on low.

If you are out of Opus credits, you are out of all model credits.
Post reply on HN