Live data from Hacker News

Claude Sonnet 5

anthropic.com

721–730 of 822 posts

Re: Claude Sonnet 5

#721

Earlier quoted context omitted.

They're actively trying to use lobbying power to make open weight models illegal. So I'm just not going to use their services at all anymore. I don't think they're a net gain if you're a skilled senior, and the hidden cost in terms of technical debt and skill atrophy is just being swept under the rug. I'll be okay without their bullshit generator.

> I don't think they're a net gain if you're a skilled senior I'm a skilled senior (I'm 54 and been coding since I was about 8; I've been 100% AI-generated code for at least 6 months now and have produced a combination of speed and quality that has astonished me; my velocity is apparent at https://github.com/pmarreck/ ) and this has been a massive net gain, so your claim is now officially in sheer defiance of reality…

> MFIC

Getting another agent to validate the first agent is a tower of sand.

> my velocity is apparent at https://github.com/pmarreck/)

Forgive me, but the active repos all look like reimplementations of existing good open source code (which of course is ideal training data) - rm_safe has rip for example. Or prototypes. Is there anything that actually has a user base > 1?

Re: Claude Sonnet 5

#722

Earlier quoted context omitted.

Like I said, Anthropic's marketing is killing it, they've got people freely(?) shilling for them on public forums so even if they have shit developer relations and community relations and a model that's mostly worse while being more expensive, they can ride a wave of misinformation.

> they have shit developer relations Not true > model that's mostly worse while being more expensive Not true > they can ride a wave of misinformation. Not true

Look at the way that Anthropic has legally threatened people who do stuff they don't like around Claude Code and their subs, and compare that to how OpenAI has acted. Look at how mixed up and unstable their communication is on policies is relative to OpenAI. Don't take my word for it, Theo/Primeagen have a whole back catalog outlining how shitty Anthropic is.

Look at the cost per intelligence of Opus vs GPT 5.5.

Anthropic is the Taylor Swift of frontier labs... Not bad, but massively, MASSIVELY stan'd for inexplicable reasons, in violation with reality.

Re: Claude Sonnet 5

#724
post #79

Wow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters. From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable than Opus 4.8 and Mythos 5 As with the other evaluations in this section, these results were achieved with all safeguards turned off. When run with our default mitigations, Sonnet 5 scored a 0 on CyberG…

Not to single you out, parent commenter, but I really hope the quality of discourse on HN will move past these basic comparisons eventually. It seems like every thread on every model release has the exact same comments. "Wow, X models is Y% better or worse than Claude Z model on T benchmark" "That's irrelevant, they're just benchmaxing." "Not useable for daily coding or agentic workloads, the vibes are totally wrong.…

I feel the same way sometimes.

I read a comment earlier that said "I think it's likely that they've scraped all the code regardless of license and trained on it, given how much they scrape the web."

That's what every other comment said like 3 years ago. Where has this guy been?

The trends in discussion about LLMs gets very, very tired--there's little added but personal opinions.

Re: Claude Sonnet 5

#725
> Me: What was the sushi place near latitude 41 in Columbus? Did it go out of business I don't see it on Google maps anymore

> Sonnet 5 (medium): None of these past chats mention a sushi place — I don't have anything on record about that. Do you remember the name, or roughly which part of Columbus (neighborhood/street) it was near? That'll help me search and check its current status.

Not impressed. It got the name right on high effort one shot but hallucinated the date relativity (Jan 2026 is not last month...). Worked okay on extra. Sonnet 4.6 worked fine on medium, high, and extra one shot.

Edit: as the replies point out, the prompt is definitely ambiguous however Sonnet 5 didn't even extract the semantic meaning "looking for near " which all the human replies seem to understand. Even Haiku 4.5 identifies the semantic meaning although it fails to retrieve the correct results

> Haiku 4.5 (reasoning off) I'll help you find information about that sushi place. Let me search for sushi restaurants that were near latitude 41 in Columbus and check their current status.

>_Searched the web_

> I see that "Latitude 41" is a restaurant in Columbus, but it's actually a modern American restaurant, not specifically a sushi place. However, based on your mention of latitude 41, that's what came up. Let me search more specifically for sushi restaurants that may have closed in Columbus around that area.

Re: Claude Sonnet 5

#726
post #79

Earlier quoted context omitted.

Not to single you out, parent commenter, but I really hope the quality of discourse on HN will move past these basic comparisons eventually. It seems like every thread on every model release has the exact same comments. "Wow, X models is Y% better or worse than Claude Z model on T benchmark" "That's irrelevant, they're just benchmaxing." "Not useable for daily coding or agentic workloads, the vibes are totally wrong.…

"It's totally obvious they quantitized Claude Z"

At least we quit with the "i asked it this question and here's what it said" comments. They were truly awful for the first 6 months or so.

Or the "I have my own personal benchmark..."

"Claude and its political bias thinks the supreme court should..."

Re: Claude Sonnet 5

#727
post #725

> Me: What was the sushi place near latitude 41 in Columbus? Did it go out of business I don't see it on Google maps anymore > Sonnet 5 (medium): None of these past chats mention a sushi place — I don't have anything on record about that. Do you remember the name, or roughly which part of Columbus (neighborhood/street) it was near? That'll help me search and check its current status. Not impressed. It got the name ri…

Try it 25 more times and let us know how it averages out. It's non-deterministic, remember?

Re: Claude Sonnet 5

#728
post #725

> Me: What was the sushi place near latitude 41 in Columbus? Did it go out of business I don't see it on Google maps anymore > Sonnet 5 (medium): None of these past chats mention a sushi place — I don't have anything on record about that. Do you remember the name, or roughly which part of Columbus (neighborhood/street) it was near? That'll help me search and check its current status. Not impressed. It got the name ri…

What was your expectation? That your prompt would trigger a web search, first, before the introspection of past conversations and a training set recall?

How did Sonnet 4.6 respond that was objectively better for your use case?

Re: Claude Sonnet 5

#729
post #727
post #725

> Me: What was the sushi place near latitude 41 in Columbus? Did it go out of business I don't see it on Google maps anymore > Sonnet 5 (medium): None of these past chats mention a sushi place — I don't have anything on record about that. Do you remember the name, or roughly which part of Columbus (neighborhood/street) it was near? That'll help me search and check its current status. Not impressed. It got the name ri…

Try it 25 more times and let us know how it averages out. It's non-deterministic, remember?

I tried 3 more times. Two were nearly identical and 1 recognized Latitude 41 as a restaurant but had a similar useless reply

Re: Claude Sonnet 5

#730
post #728
post #725

> Me: What was the sushi place near latitude 41 in Columbus? Did it go out of business I don't see it on Google maps anymore > Sonnet 5 (medium): None of these past chats mention a sushi place — I don't have anything on record about that. Do you remember the name, or roughly which part of Columbus (neighborhood/street) it was near? That'll help me search and check its current status. Not impressed. It got the name ri…

What was your expectation? That your prompt would trigger a web search, first, before the introspection of past conversations and a training set recall? How did Sonnet 4.6 respond that was objectively better for your use case?

[flagged]
Post reply on HN