Earlier quoted context omitted.
There's two classes of models now - the cybersecurity ones that none of us are getting, and the 'safe' models released for general consumption. This is letting us know which side of the divide it sits on.
this seems rather counter-productive, wouldn't a model with less cybersecurity capabilities be more likely to produce insecure code? Not to mention, Chinese models don't have these restrictions and can be used to exploit said unsecure code. I supposed I shouldn't be surprised at how the trump admin is approaching AI regulation, counter-productive is really all they do
Claude Sonnet 5
351–360 of 822 posts
Re: Claude Sonnet 5
#352Earlier quoted context omitted.
> I have been moving more and more to K2.7 Code and GLM-5.2 the last few weeks. They are often good enough for assistance, very fast, and cheap. I've moved completely to local models that I run with my M1 Mac Studio (64gb ram) some time ago. But for the rare times when I feel the local, quantized Qwen3.6 isn't enough, I just connect to Openrouter and use something like Kimi, GLM or Deepseek for a fraction of the pric…
Which quant do you use? I have a similar setup and the speed is atrocious at 4-bit.
https://huggingface.co/mlx-community/Qwen3.6-35B-A3B-OptiQ-4...
Re: Claude Sonnet 5
#353Earlier quoted context omitted.
No it doesn't? It's worse than Opus across the whole shared frontier on both plots.
Agreed. The graphs clearly show that opus 4.8 performs strictly better at the same cost per task
The graphs show parts of the cost/performance pareto frontier occupied by Opus 4.8 and others occupied by Sonnet 5.0. If Opus 4.8 was strictly better at cost per task like you say, by definition the entire frontier would be occupied by Opus.
So neither is pareto-dominant over the other. In contrast, Sonnet 5.0 is Pareto-dominent over Sonnet 4.6 on those graphs.
Re: Claude Sonnet 5
#354Re: Claude Sonnet 5
#355Earlier quoted context omitted.
I am deeply surprised by the silence of philosophers, sociologists, liberal arts majors, economists. Where are the think tanks who contemplate and debate the societal aspects? The tech is advancing full steam but the "other side" doesn't feel anywhere nearly ready.
Idk why you're perceiving silence. Feels to me like this is the main thing people talk about nowadays.
The frontier labs, on the other hand, are thinking about replacing all human labor, ending death, and the risk of it causing human extinction. Most of the apparatus we're talking about approach it very parochially; it's almost like they're embarrassed to take the grander ideas even a little seriously, for being too nerdy/sci-fi.
Re: Claude Sonnet 5
#356The cost per task chart is telling me that I should _never_ use Sonnet 5 above medium effort level - Opus always performs better for a given cost. So I guess the takeaway is that if Sonnet 5 medium isn't good enough for you, switch models, not effort levels.
Re: Claude Sonnet 5
#357Wow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters. From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable than Opus 4.8 and Mythos 5 As with the other evaluations in this section, these results were achieved with all safeguards turned off. When run with our default mitigations, Sonnet 5 scored a 0 on CyberG…
Not to single you out, parent commenter, but I really hope the quality of discourse on HN will move past these basic comparisons eventually. It seems like every thread on every model release has the exact same comments. "Wow, X models is Y% better or worse than Claude Z model on T benchmark" "That's irrelevant, they're just benchmaxing." "Not useable for daily coding or agentic workloads, the vibes are totally wrong.…
I generally agree with this in spirit https://www.seangoedecke.com/are-new-models-good/ , but I think you can read Anthropic's results showing Sonnet 5 as almost strictly worse than Opus 4.8 as very credible/meaningful, and then draw comparisons from that
Re: Claude Sonnet 5
#358interesting how much worse the sentiment around Anthropic is getting
Seems like a combination of multiple factors: "They took my shit away!" -- 3-day Fable 5 addicts (me) "How dare they tell Trump no?" -- US nationalist / "my country right or wrong" types "Great to see a closed source company fail!" -- open source boosters "Great to see an American company fail!" -- anti-US, and/or pro-China folks "Great to see a successful company fail!" -- anti-capitalists and/or sour-grapes crab bu…
Re: Claude Sonnet 5
#359Re: Claude Sonnet 5
#360Earlier quoted context omitted.
> This is why there was a "Sonnet only" usage bar for Max tier for the longest time. it's still there. I still don't totally grok why I can't use all my tokens on Sonnet if I want to... maybe that signals something?
They want to encourage diversifying model use.