Live data from Hacker News

Claude Sonnet 5

anthropic.com

781–790 of 822 posts

Re: Claude Sonnet 5

#781
post #480

Earlier quoted context omitted.

That's possibly the worst pelican I saw from all recent LLMs. Meanwhile GLM 5.2 drew a cool self-contained fully animated SVG pelican: https://simonwillison.net/2026/Jun/17/glm-52

Just need the legs to interact with the pedals now :D

That's something that Sonnet 5 couldn't get right even on a static SVG.

Re: Claude Sonnet 5

#783
post #635

Earlier quoted context omitted.

Have you really found claude to much more more capable than eg deepseek? Anthropic has little to no chance of producing a competitive business model in the long term.

The cheap models are cost-competitive if you are running them in long-running agentive tasks. But they take a lot longer to reach the same goal for complex tasks, so the difference is still very real, and the cost-savings are still very much a question of how well you manage to characterise the tasks they will do quickly and pick and choose what to use when. I kind of agree that I think the cheap models will eat away…

[deleted]

Re: Claude Sonnet 5

#784
post #457

Earlier quoted context omitted.

Your benchmark has Gemini 3.5 Flash as the best model, which doesn't compute for me

It is on top for many benchmarks, only not the coding/agentic ones. Still one of the most intelligent models overall, most likely to get any question you ask correctly (without tools).

Not in my experience, it tends to pick up subtle orientations given in a question (like "which is better, A and B?" and in the context you add you list a few things for A and B) and will absolutely run with them even if they're not true. Has been an issue with Gemini models since at least 3.0. Maybe that makes them great roleplaying models, but for factual information they just run with the slightest hint in one direction or another and never really push back objectively.

Re: Claude Sonnet 5

#785

I'm struggling to understand why I'd ever use this instead of just using a lower effort level for opus given on many of the benchmarks listed the cost per task rises above opus at anything higher than medium effort. Only thing I can think of is for when someone is out of opus credits. Of course there are API billing use cases but I'd probably still just use opus on low.

Looking at some of the agentic coding benchmarks on the system card[0], pages 117-118, it seems that running it at low outperforms Sonnet 4.6 at any level, and is a good deal cheaper as well. So on low it could be a good workhorse for an Opus-planned task. [0] https://www.anthropic.com/claude-sonnet-5-system-card

That is certainly an improvement then. Sonnet 4.6 is a great everyday agent for the limited Pro plan, but it’s not much better than M3 or Kimi 2.7, both significantly cheaper models.

Re: Claude Sonnet 5

#786

Earlier quoted context omitted.

What you are doing, is producing an unnecessary summary of the result, not reasoning that models do to come up with the result. I don't get what value you get out of this.

This is the same concept as Chain of Thought. Just that when "native" thinking is used it's hidden from the end result via special tags. If you force a model to reason about a result before producing the result, you get more accurate results. Because "reason" comes before the selection, it has to think through why it is producing the result beforehand (e.g. produce a block of text that makes the correct answer statis…

I'm aware how reasoning works. But maybe I am misunderstanding how you use it.

Because from what I see, you are not getting a chain of thought, but just a summary of a non-chain-of-thought answer, with no reasoning either explicit or implicit.

Re: Claude Sonnet 5

#787

Earlier quoted context omitted.

Mini nukes are hard to build. They require the entire industrial base to produce. The knowledge how to build them is universally available. The model can tell you how to refine weapon-grade plutonium, but will not get you a factory for that - and it is genuinely hard to build. What you are describing is pure information control. Which is supposed to be operated by the people who are the most ill equipped to do that -…

Superintelligence itself is the mini nuke. With it, you can hack virtually any system, order some viral precursors and put together a bioweapon at home (yes you can actually do this, they are available for sale freely), etc.

Not exactly. I used to run my home biolab for CRISPR experiments (thanks to TheOdin for hardware and materials). This is harder than you think. Cultivating viruses is hard. Bioweapons are not about synthesizing new viruses (though this alone is a hard problem - you need growth environments and a lot of lab hardware) - they are about weaponization, aerosol production, hardening tests. You'll need to spend a few millions just for the prototype stage. And if you are into bioweaponeering, you'd better source some actual bio experts, they are quite available.

Knowledge is widely available already. Regulations there are not about witholding the knowledge.

Re: Claude Sonnet 5

#788

Earlier quoted context omitted.

> I have been moving more and more to K2.7 Code and GLM-5.2 the last few weeks. They are often good enough for assistance, very fast, and cheap. I've moved completely to local models that I run with my M1 Mac Studio (64gb ram) some time ago. But for the rare times when I feel the local, quantized Qwen3.6 isn't enough, I just connect to Openrouter and use something like Kimi, GLM or Deepseek for a fraction of the pric…

What is your motivation? Privacy and/or data protection? I currently don't see a world where it makes sense to run a local model that will eats up 60% of my RAM, 20-30% of my disk space while providing worse quality output than a $20/month subscription.

> What is your motivation? Privacy and/or data protection?

For me, those things are nice benefits but not the motivation. It's just a desire to own my tools and remove the magic wherever possible, and not pay a big corporation who is constantly tweaking/changing things that I might not like. It's the same reason I migrated everything I host off of Azure and onto a VPS, and why I moved from Jetbrains Rider to Neovim, and so on.

Re: Claude Sonnet 5

#789
post #132

Wow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters. From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable than Opus 4.8 and Mythos 5 As with the other evaluations in this section, these results were achieved with all safeguards turned off. When run with our default mitigations, Sonnet 5 scored a 0 on CyberG…

I have tried to rewrite an article with GLM-5.2 and with Sonnet 4.6. Completely different results as LLM is non-deterministic. But GLM-5.2 made a lot of subtle mistakes that needed to be corrected by hand. On the opposite, Sonnet found and corrected all mistakes in the second round. Similar situation was with planning and coding. GLM-5.2 seems to be good “on paper” but the real usage results was different. And I am n…

I’m just dipping my feet in the water of local models and I really feel this. I had a simple alignment task (align known quality transcript without timestamps with timestamped but lower quality whisper transcriptions) and I went through 12 rounds of testing across 4 generations of 3 models. The results were all over the map even across versions of the same model. Spin out the task to something as big as coding and wow.

If you have any advice/blogs on doing project specific benchmarks I’d love to hear it. I’m trying, but it’s haphazard at the moment

Re: Claude Sonnet 5

#790
post #608

Earlier quoted context omitted.

> Anthropic has little to no chance of producing a competitive business model in the long term. Extraordinary thing to say about the fastest growing company in the history of capitalism. They will soon have access to public markets, essentially unlimited capital, and can build insanely large models that they don't have to make public... ever. They can just use those models to run their business, train better models,…

> Extraordinary thing to say about the fastest growing company in the history of capitalism. They will soon have access to public markets, essentially unlimited capital There is no such thing as unlimited capital. The faster they grow the faster they burn capital. Eventually it will run out.

That's not how economies work. It's not a fixed pie that everyone shares.
Post reply on HN