Live data from Hacker News

Claude Sonnet 5

anthropic.com

361–370 of 822 posts

Re: Claude Sonnet 5

#361
post #332

Earlier quoted context omitted.

> This is why there was a "Sonnet only" usage bar for Max tier for the longest time. it's still there. I still don't totally grok why I can't use all my tokens on Sonnet if I want to... maybe that signals something?

They want to encourage diversifying model use.

Why?

Re: Claude Sonnet 5

#362

Earlier quoted context omitted.

What I want is a harness that knows how to optimize this kind of thing for me.

You might want to check out Amp: https://ampcode.com/

I appreciate the suggestion! But it isn't clear to me, from reading their marketing site, what they bring to the table from this perspective. Can you give me a more targeted pitch?

Re: Claude Sonnet 5

#363

Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models. I have been using Sonnet 4.6 more than Opus, because I'm mostly doing agent-assisted development and not fully agent-driven development. This announcement does not make me positive, I have fou…

I've been saying for ages that since Opus 4.6 models are increasingly smarter but further unhelpful as assistants. Fable was amazing as a vibecoder but as an assistant it can't resist jumping into implementation and filling chats of pointless jargon. It's really grim if you're looking for assistance instead of an implementor. GPT 5.5 Pro and Fable are gorgeous bullshitters that pretend to be right (often convincingly…

Just to follow up on what I mean, this was my first interaction with Sonnet 5:

"I just cloned this repo, investigate how to set it up, don't install anything, just collect information"

_spews information_

I proceed with the setup, but get a Linux specific dependency in a bash script, so I want to evaluate whether it can be rewritten...

"There's this error on MacOS, I think it's because we need linux-utils from brew, verify whether the script can be written in bare posix"

_proceeds installing linux-utils and all the rest_

"Didn't I tell you to not install anything?"

_you're absolutely right_

F*k me..

Re: Claude Sonnet 5

#364
post #157

Earlier quoted context omitted.

Due to Dario hyping it up as a world ending model. If they kept their mouths shut we'd all have it now still.

Where is gpt 5.6?

If not for Dario hyping Mythos and Fable, GPT 5.6 would've released just fine on schedule as a point release without all the fear mongering. It was because Fable was banned that now the government is scrutinizing all models.

Re: Claude Sonnet 5

#365

I'm struggling to understand why I'd ever use this instead of just using a lower effort level for opus given on many of the benchmarks listed the cost per task rises above opus at anything higher than medium effort. Only thing I can think of is for when someone is out of opus credits. Of course there are API billing use cases but I'd probably still just use opus on low.

More and more I find myself trying to stop Opus from doing something stupid, and at every turn I need to tell it to stop overcomplicating things. I think the models are being optimized for wealth extraction from users and companies, instead of solving problems. I don't know why Opus would try to create an entire library when I told it specifically to do something simple that would take 2-3 lines of Python.

> More and more I find myself trying to stop Opus from doing something stupid, and at every turn I need to tell it to stop overcomplicating things

Yeah, that’s my thoughts as well. I feel it’s great for benchmarks and some tasks while in other it tries to spend as much tokens as possible, tries to overcomplicate task and needs seconds or third round of steering that costs. With the scale Anthropic operates I bet it’s huge amount of extra money just to make sure their model works.

Re: Claude Sonnet 5

#366

Earlier quoted context omitted.

There’s no way to justify their valuations if they get downgraded to a pair programming tool. They need fully agentic stuff to work and replace human engineers to even come close. Offhand, I’m not even certain whether a model like that could justify the constant retraining we’re doing on the agentic models. It doesn’t make a lot of sense to spend millions or billions on training to reduce hallucinations by 0.3% if yo…

> no way to justify their valuations if they get downgraded to a pair programming tool I think there is. Pair today doesn’t mean they’re locked into that forever.

Their valuations don't make sense as just programming tools, period. Forget about if they are still human driven.

Re: Claude Sonnet 5

#367
post #276

Earlier quoted context omitted.

There’s no way to justify their valuations if they get downgraded to a pair programming tool. They need fully agentic stuff to work and replace human engineers to even come close. Offhand, I’m not even certain whether a model like that could justify the constant retraining we’re doing on the agentic models. It doesn’t make a lot of sense to spend millions or billions on training to reduce hallucinations by 0.3% if yo…

Dario has publicly claimed each model has been profitable, even accounting for its training costs; it's just that each new model is exponentially more expensive to train than the last, so the income lags and it looks like the company is losing money overall. Now, we can't know if this is true unfortunately, but it's not directly contradicted by anything that's known publicly at least. I thought it was an interesting…

A common extreme misconception is that inference is expensive and that providers are loosing a lot of money. Inference is extremely lucrative and profitable.

Re: Claude Sonnet 5

#368
Let’s see how long until opus 5 comes out but to me this lends some credence to the rumour that fable/mythos was supposed to be opus 5

Re: Claude Sonnet 5

#369
post #353
post #339

Earlier quoted context omitted.

Agreed. The graphs clearly show that opus 4.8 performs strictly better at the same cost per task

But they don't show "strictly better" performance at cost per task! The graphs show parts of the cost/performance pareto frontier occupied by Opus 4.8 and others occupied by Sonnet 5.0. If Opus 4.8 was strictly better at cost per task like you say, by definition the entire frontier would be occupied by Opus. So neither is pareto-dominant over the other. In contrast, Sonnet 5.0 is Pareto-dominent over Sonnet 4.6 on th…

> by definition the entire frontier would be occupied by Opus.

But the entire frontier is occupied by Opus under any reasonable interpolation scheme (piecewise linear which is what they've done, and most reasonable spline or polynomial fits would also lead to the same result) over the overlapping x values for which both are defined.

Under that interpolation scheme, for x > ($ cost of Opus low effort), Opus is Pareto-dominant over Sonnet 5. You can see this by picking any point on Opus's interpolation and realizing that you get strictly worse by switching to Sonnet for the same x value or the same y value. Meaning if you want to pay the same $x then you get a worse y, or if you want the same y you pay more $x.

Re: Claude Sonnet 5

#370
post #239

Earlier quoted context omitted.

I find these nefarious intention theories shallow. It can both be the case that the endstate is them owning the means of production without that being the intended guiding goal. Companies can chase profit without being Leninistic boogeymen.

There is no nefariousness in owning all the means of production, it's the endgame of maximizing profit. However the result is exactly the same, concentration of power.

No nefariousness other than the subjugation of the majority of humanity? You're insane
Post reply on HN