Live data from Hacker News

Claude Sonnet 5

anthropic.com

391–400 of 822 posts

Re: Claude Sonnet 5

#391

The cost per task chart is telling me that I should _never_ use Sonnet 5 above medium effort level - Opus always performs better for a given cost. So I guess the takeaway is that if Sonnet 5 medium isn't good enough for you, switch models, not effort levels.

You're referring to the Agentic search, but if you look at the Agentic computer use the cost is basically halved. However, I am also confused about market positioning. Too expensive to perform daily tasks - open souce models are much cheaper - and not frontier model to address complex real world problems. Rarely used Sonnet btw.

> Too expensive to perform daily tasks - open souce models are much cheaper

There is a real advantage, especially for businesses, in using an off the shelf solution from a corporate provider.

Personally, the advantage of not having to set up multiple solutions from multiple sources outweighs the cost of a $20 a month subscription. Think about why a lot of consumers prefer Apple devices over Linux. There are a lot of advantages to Linux, but "never having to think about my tools" is its own advantage.

Re: Claude Sonnet 5

#392
post #162
post #89

Earlier quoted context omitted.

There's also Chinese models, which aren't trying to self-limit capabilities.

…as long as you don’t ask them about certain dates or squares. Also, I wouldn’t expect Mythos-class models to be allowed to be openly released by the CCP. Thinking otherwise is pure naivety.

Like the sibling said, you can fine tune if the rejections are in the weights but most often it's actually in the API harness itself; download Qwen or DeepSeek and run it locally to ask about certain dates and squares and it will happily tell you.

Re: Claude Sonnet 5

#393

Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models. I have been using Sonnet 4.6 more than Opus, because I'm mostly doing agent-assisted development and not fully agent-driven development. This announcement does not make me positive, I have fou…

gemma-4-e4b is very good at assistance too, and is local and fast and small (and "free")

Re: Claude Sonnet 5

#394
post #334

Earlier quoted context omitted.

The world is not zero sum. Value is created, not just preserved. Anthropic and OpenAI creating value does not imply that smaller guys can not also create value.

But marketplaces also exist and big players in a marketplace are often able to manipulate the market such that they are advantaged and small players are not able to break in.

This is true of every market that has ever existed, and that's not stopped small players from finding niches.

Re: Claude Sonnet 5

#395
post #388

I just tested it on my benchmarks[0], it's GLM-5.2 level, at 2x cost, but also 2x faster. Weak spots (categories it fails): - Trivia — 0/3 - basically not much built-in knowledge - Combined tool-calling tasks — score 45/100, sometimes makes invalid tool calls - Puzzle Solving — score 77, flubs carwash-like tests [0]: https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-med...

As always, note: faster than GLM-5.2 doesn't mean too much, as GLM-5.2 is served by different providers, so the inference speed can vary drastically between providers or over time.

Re: Claude Sonnet 5

#396
post #68

Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models. I have been using Sonnet 4.6 more than Opus, because I'm mostly doing agent-assisted development and not fully agent-driven development. This announcement does not make me positive, I have fou…

I've been largely disappointed how much the Claude models ignore custom instructions, and sometimes even prompts on the chat interface. It sometimes feels like talking to a wall, or as if there was a third person in the chatroom whose messages I can't see. I can't help but feel this is intentional towards the 'Agentic' workflow.

[dead]

Re: Claude Sonnet 5

#397

What is the reference, unbiased, honest, reputable and trustworthy site that ranks and compare models on the couple of realistic metrics that matters ? ("Does it work for code", "no, I mean, for real", "how much does it cost", etc...) ?

The only metric that worked for me is running the same prompt 5x for each LLMs on my projects.

I keep specific branches a state where they are ready to develop new features.

Re: Claude Sonnet 5

#400

The cost per task chart is telling me that I should _never_ use Sonnet 5 above medium effort level - Opus always performs better for a given cost. So I guess the takeaway is that if Sonnet 5 medium isn't good enough for you, switch models, not effort levels.

Well, it is a Sonnet model, it is indeed better[0] than Sonnet 4.6 (smarter, faster, cheaper), but I don't see why would you use it as opposed to Opus 4.8 low or GLM-5.2...

[0]: https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-med...

Post reply on HN