Earlier quoted context omitted.
That's possibly the worst pelican I saw from all recent LLMs. Meanwhile GLM 5.2 drew a cool self-contained fully animated SVG pelican: https://simonwillison.net/2026/Jun/17/glm-52
Just need the legs to interact with the pedals now :D
Claude Sonnet 5
781–790 of 822 posts
Re: Claude Sonnet 5
#782Re: Claude Sonnet 5
#783Earlier quoted context omitted.
Have you really found claude to much more more capable than eg deepseek? Anthropic has little to no chance of producing a competitive business model in the long term.
The cheap models are cost-competitive if you are running them in long-running agentive tasks. But they take a lot longer to reach the same goal for complex tasks, so the difference is still very real, and the cost-savings are still very much a question of how well you manage to characterise the tasks they will do quickly and pick and choose what to use when. I kind of agree that I think the cheap models will eat away…
Re: Claude Sonnet 5
#784Earlier quoted context omitted.
Your benchmark has Gemini 3.5 Flash as the best model, which doesn't compute for me
It is on top for many benchmarks, only not the coding/agentic ones. Still one of the most intelligent models overall, most likely to get any question you ask correctly (without tools).
Re: Claude Sonnet 5
#785I'm struggling to understand why I'd ever use this instead of just using a lower effort level for opus given on many of the benchmarks listed the cost per task rises above opus at anything higher than medium effort. Only thing I can think of is for when someone is out of opus credits. Of course there are API billing use cases but I'd probably still just use opus on low.
Looking at some of the agentic coding benchmarks on the system card[0], pages 117-118, it seems that running it at low outperforms Sonnet 4.6 at any level, and is a good deal cheaper as well. So on low it could be a good workhorse for an Opus-planned task. [0] https://www.anthropic.com/claude-sonnet-5-system-card
Re: Claude Sonnet 5
#786Earlier quoted context omitted.
What you are doing, is producing an unnecessary summary of the result, not reasoning that models do to come up with the result. I don't get what value you get out of this.
This is the same concept as Chain of Thought. Just that when "native" thinking is used it's hidden from the end result via special tags. If you force a model to reason about a result before producing the result, you get more accurate results. Because "reason" comes before the selection, it has to think through why it is producing the result beforehand (e.g. produce a block of text that makes the correct answer statis…
Because from what I see, you are not getting a chain of thought, but just a summary of a non-chain-of-thought answer, with no reasoning either explicit or implicit.
Re: Claude Sonnet 5
#787Earlier quoted context omitted.
Mini nukes are hard to build. They require the entire industrial base to produce. The knowledge how to build them is universally available. The model can tell you how to refine weapon-grade plutonium, but will not get you a factory for that - and it is genuinely hard to build. What you are describing is pure information control. Which is supposed to be operated by the people who are the most ill equipped to do that -…
Superintelligence itself is the mini nuke. With it, you can hack virtually any system, order some viral precursors and put together a bioweapon at home (yes you can actually do this, they are available for sale freely), etc.
Knowledge is widely available already. Regulations there are not about witholding the knowledge.
Re: Claude Sonnet 5
#788Earlier quoted context omitted.
> I have been moving more and more to K2.7 Code and GLM-5.2 the last few weeks. They are often good enough for assistance, very fast, and cheap. I've moved completely to local models that I run with my M1 Mac Studio (64gb ram) some time ago. But for the rare times when I feel the local, quantized Qwen3.6 isn't enough, I just connect to Openrouter and use something like Kimi, GLM or Deepseek for a fraction of the pric…
What is your motivation? Privacy and/or data protection? I currently don't see a world where it makes sense to run a local model that will eats up 60% of my RAM, 20-30% of my disk space while providing worse quality output than a $20/month subscription.
For me, those things are nice benefits but not the motivation. It's just a desire to own my tools and remove the magic wherever possible, and not pay a big corporation who is constantly tweaking/changing things that I might not like. It's the same reason I migrated everything I host off of Azure and onto a VPS, and why I moved from Jetbrains Rider to Neovim, and so on.
Re: Claude Sonnet 5
#789Wow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters. From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable than Opus 4.8 and Mythos 5 As with the other evaluations in this section, these results were achieved with all safeguards turned off. When run with our default mitigations, Sonnet 5 scored a 0 on CyberG…
I have tried to rewrite an article with GLM-5.2 and with Sonnet 4.6. Completely different results as LLM is non-deterministic. But GLM-5.2 made a lot of subtle mistakes that needed to be corrected by hand. On the opposite, Sonnet found and corrected all mistakes in the second round. Similar situation was with planning and coding. GLM-5.2 seems to be good “on paper” but the real usage results was different. And I am n…
If you have any advice/blogs on doing project specific benchmarks I’d love to hear it. I’m trying, but it’s haphazard at the moment
Re: Claude Sonnet 5
#790Earlier quoted context omitted.
> Anthropic has little to no chance of producing a competitive business model in the long term. Extraordinary thing to say about the fastest growing company in the history of capitalism. They will soon have access to public markets, essentially unlimited capital, and can build insanely large models that they don't have to make public... ever. They can just use those models to run their business, train better models,…
> Extraordinary thing to say about the fastest growing company in the history of capitalism. They will soon have access to public markets, essentially unlimited capital There is no such thing as unlimited capital. The faster they grow the faster they burn capital. Eventually it will run out.