Live data from Hacker News

Claude Sonnet 5

anthropic.com

171–180 of 822 posts

Re: Claude Sonnet 5

#171

Earlier quoted context omitted.

The reality is that Fable will eventually be obsolete and Sonnet / Opus will surpass it. Fable did cost 2x as much as Opus, so I assume it involves a much higher cost for what it did, but I wouldn't be surprised if Fable will be obsoleted by Opus or even Sonnet sooner or later at less cost.

Okay I don’t care about “eventually”, I want Fable now.

Have you considered getting better at coding so you can build stuff yourself instead of waiting for models you might not be able to get access to anymore?

Re: Claude Sonnet 5

#172

Earlier quoted context omitted.

"Lower ability to perform cybersecurity-related tasks" makes me super concerned it will leave my codebase like Swiss cheese for any American granny with access to Fable 5, when we non-American Brits, or rest-of-worlders, don't have access to it to clean our codebases.

I think they don’t understand that cybersecurity skills are what prevent bad code from making it into production. It’s like telling a chef to cook without a knife because knives can kill people. Dario and his lackeys at Anthropic aren’t visionaries.

I think this is more aimed at the US gov't than anything. They want to be clear that it's not very good at hacking, so that the gov't won't ban it.

I'm sure they're well-aware that this also will make it worse at building secure systems, but the gov't isn't restricting releases based on that.

Re: Claude Sonnet 5

#173

The cost per task chart is telling me that I should _never_ use Sonnet 5 above medium effort level - Opus always performs better for a given cost. So I guess the takeaway is that if Sonnet 5 medium isn't good enough for you, switch models, not effort levels.

There are two wrinkles to this: - For Claude.ai subscriptions I think Sonnet is much cheaper than Opus. This is why there was a "Sonnet only" usage bar for Max tier for the longest time. - For some tasks the sheer amount of raw input tokens is the most important. For example multimodal computer use tasks. You can't make them any more efficient on Opus by turning down the reasoning, so a cheaper model like Sonnet is u…

> This is why there was a "Sonnet only" usage bar for Max tier for the longest time.

it's still there. I still don't totally grok why I can't use all my tokens on Sonnet if I want to... maybe that signals something?

Re: Claude Sonnet 5

#174
Why is Claude Sonnet 5 allowed to be released but OpenAI Terra not? Are they not the same class of models?

Re: Claude Sonnet 5

#175

Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models. I have been using Sonnet 4.6 more than Opus, because I'm mostly doing agent-assisted development and not fully agent-driven development. This announcement does not make me positive, I have fou…

agent-assisted development uses orders of magnitude fewer tokens than agent-driven development

the incentives aren't there sadly

Re: Claude Sonnet 5

#176

Earlier quoted context omitted.

Yeah, there's a real opportunity for one of these companies to invest time in a model that's tuned for, to use your term, agent-assisted developement. Trouble is, everyone inside their buildings seems to believe that no one will be working like that in a year or two.

There’s no way to justify their valuations if they get downgraded to a pair programming tool. They need fully agentic stuff to work and replace human engineers to even come close. Offhand, I’m not even certain whether a model like that could justify the constant retraining we’re doing on the agentic models. It doesn’t make a lot of sense to spend millions or billions on training to reduce hallucinations by 0.3% if yo…

> no way to justify their valuations if they get downgraded to a pair programming tool

I think there is. Pair today doesn’t mean they’re locked into that forever.

Re: Claude Sonnet 5

#177
post #14

I didn't think they'd actually release a model that was worse than the open-weight frontier and at a higher price-point. Wow.

That's yet to be determined. I think a lot of open-weight models are benchmaxxed and their usefulness for many tasks are not represented by those.

Re: Claude Sonnet 5

#179
post #162
post #89

Earlier quoted context omitted.

There's also Chinese models, which aren't trying to self-limit capabilities.

…as long as you don’t ask them about certain dates or squares. Also, I wouldn’t expect Mythos-class models to be allowed to be openly released by the CCP. Thinking otherwise is pure naivety.

Well, the weights are open. De-CCP-ing them is a trivial task, about 40 minutes on modern hardware. So can be done for about $50.
Post reply on HN