Live data from Hacker News

Claude Sonnet 5

anthropic.com

351–360 of 822 posts

Re: Claude Sonnet 5

#351
post #93

Earlier quoted context omitted.

There's two classes of models now - the cybersecurity ones that none of us are getting, and the 'safe' models released for general consumption. This is letting us know which side of the divide it sits on.

this seems rather counter-productive, wouldn't a model with less cybersecurity capabilities be more likely to produce insecure code? Not to mention, Chinese models don't have these restrictions and can be used to exploit said unsecure code. I supposed I shouldn't be surprised at how the trump admin is approaching AI regulation, counter-productive is really all they do

As contradictory as it sounds, they (Anthropic) are probably trying to dance the fine line where its public models can write secure code but cannot exploit insecure code.

Re: Claude Sonnet 5

#352

Earlier quoted context omitted.

> I have been moving more and more to K2.7 Code and GLM-5.2 the last few weeks. They are often good enough for assistance, very fast, and cheap. I've moved completely to local models that I run with my M1 Mac Studio (64gb ram) some time ago. But for the rare times when I feel the local, quantized Qwen3.6 isn't enough, I just connect to Openrouter and use something like Kimi, GLM or Deepseek for a fraction of the pric…

Which quant do you use? I have a similar setup and the speed is atrocious at 4-bit.

I'm using 4-bit as well, with the MoE model. I also use the MLX versions which are optimized for Apple CPUs (from what I understand anyway, I'm just an LLM layman). According to my oMLX dashboard, I'm getting about 50 tokens per second out of this model – not blazing fast, but more than fast enough to be useful to me.

https://huggingface.co/mlx-community/Qwen3.6-35B-A3B-OptiQ-4...

Re: Claude Sonnet 5

#353
post #339

Earlier quoted context omitted.

No it doesn't? It's worse than Opus across the whole shared frontier on both plots.

Agreed. The graphs clearly show that opus 4.8 performs strictly better at the same cost per task

But they don't show "strictly better" performance at cost per task!

The graphs show parts of the cost/performance pareto frontier occupied by Opus 4.8 and others occupied by Sonnet 5.0. If Opus 4.8 was strictly better at cost per task like you say, by definition the entire frontier would be occupied by Opus.

So neither is pareto-dominant over the other. In contrast, Sonnet 5.0 is Pareto-dominent over Sonnet 4.6 on those graphs.

Re: Claude Sonnet 5

#355
post #313

Earlier quoted context omitted.

I am deeply surprised by the silence of philosophers, sociologists, liberal arts majors, economists. Where are the think tanks who contemplate and debate the societal aspects? The tech is advancing full steam but the "other side" doesn't feel anywhere nearly ready.

Idk why you're perceiving silence. Feels to me like this is the main thing people talk about nowadays.

It has to do with the scope of what they're discussing. It seems extraordinarily small: e.g. what if AI increases productivity growth by 0.4%? Do data centers use too much water? Are AIs racist when reviewing resumes?

The frontier labs, on the other hand, are thinking about replacing all human labor, ending death, and the risk of it causing human extinction. Most of the apparatus we're talking about approach it very parochially; it's almost like they're embarrassed to take the grander ideas even a little seriously, for being too nerdy/sci-fi.

Re: Claude Sonnet 5

#356

The cost per task chart is telling me that I should _never_ use Sonnet 5 above medium effort level - Opus always performs better for a given cost. So I guess the takeaway is that if Sonnet 5 medium isn't good enough for you, switch models, not effort levels.

It's funny the exact same thing happened to Gemini 3.5 flash. Cheaper and more agentic model that ends up worse and more expensive than 3.5 pro low.

Re: Claude Sonnet 5

#357
post #79

Wow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters. From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable than Opus 4.8 and Mythos 5 As with the other evaluations in this section, these results were achieved with all safeguards turned off. When run with our default mitigations, Sonnet 5 scored a 0 on CyberG…

Not to single you out, parent commenter, but I really hope the quality of discourse on HN will move past these basic comparisons eventually. It seems like every thread on every model release has the exact same comments. "Wow, X models is Y% better or worse than Claude Z model on T benchmark" "That's irrelevant, they're just benchmaxing." "Not useable for daily coding or agentic workloads, the vibes are totally wrong.…

Yeah you definitely have to be skeptical regarding sentiment for open/local model capabilities, since there's bias from what people want to be true.

I generally agree with this in spirit https://www.seangoedecke.com/are-new-models-good/ , but I think you can read Anthropic's results showing Sonnet 5 as almost strictly worse than Opus 4.8 as very credible/meaningful, and then draw comparisons from that

Re: Claude Sonnet 5

#358

interesting how much worse the sentiment around Anthropic is getting

Seems like a combination of multiple factors: "They took my shit away!" -- 3-day Fable 5 addicts (me) "How dare they tell Trump no?" -- US nationalist / "my country right or wrong" types "Great to see a closed source company fail!" -- open source boosters "Great to see an American company fail!" -- anti-US, and/or pro-China folks "Great to see a successful company fail!" -- anti-capitalists and/or sour-grapes crab bu…

Most of these are good points though with the right framing.

Re: Claude Sonnet 5

#359
In my case, 4.6 degraded massively over time. 5 fails the same basic tasks that I gave 4.6 yesterday. And quite frankly this low, med, high, extra, max, turbo, ultra, ludicrous nonsense is getting tiresome

Re: Claude Sonnet 5

#360
post #332

Earlier quoted context omitted.

> This is why there was a "Sonnet only" usage bar for Max tier for the longest time. it's still there. I still don't totally grok why I can't use all my tokens on Sonnet if I want to... maybe that signals something?

They want to encourage diversifying model use.

Seems kinda weird - it's cognitive load I'd love to avoid. If I'm going to take it on, I might as well try other providers.
Post reply on HN