Live data from Hacker News

Claude Sonnet 5

anthropic.com

281–290 of 822 posts

Re: Claude Sonnet 5

#281

Earlier quoted context omitted.

You're the second person that has said this but I cannot understand why you are interpreting the "Agentic computer use" graph in this manner. The graph shows that Opus is cheaper than Sonnet for the same performance. Unless I am suffering a cognitive blindness thing right now.

Wrong! Look at it better. It shows that Opus has superior performance but at higher cost.

No, that's apples and oranges. You need to compare Sonnet5's 79% with the interpolated Opus4.8's 79%.

Re: Claude Sonnet 5

#282
post #221

Earlier quoted context omitted.

"We can raise prices in two ways: (1) raise the price per token and (2) increase the number of tokens we generate on your behalf. We promise not to do (2) maliciously. Promise."

Wouldn't it be more malicious for them not to mention this at all?

Sure, but I think doing it this way allows them to later on say they were transparent about it. Completely hiding this would make it very difficult for them excuse when getting caught.

Re: Claude Sonnet 5

#283

Earlier quoted context omitted.

You're the second person that has said this but I cannot understand why you are interpreting the "Agentic computer use" graph in this manner. The graph shows that Opus is cheaper than Sonnet for the same performance. Unless I am suffering a cognitive blindness thing right now.

Wrong! Look at it better. It shows that Opus has superior performance but at higher cost.

That is a bad comparison. Compare Sonnet xhigh against Opus medium, which is both better and cheaper.

Re: Claude Sonnet 5

#284

Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models. I have been using Sonnet 4.6 more than Opus, because I'm mostly doing agent-assisted development and not fully agent-driven development. This announcement does not make me positive, I have fou…

> I have been moving more and more to K2.7 Code and GLM-5.2 the last few weeks. They are often good enough for assistance, very fast, and cheap. I've moved completely to local models that I run with my M1 Mac Studio (64gb ram) some time ago. But for the rare times when I feel the local, quantized Qwen3.6 isn't enough, I just connect to Openrouter and use something like Kimi, GLM or Deepseek for a fraction of the pric…

Which quant do you use? I have a similar setup and the speed is atrocious at 4-bit.

Re: Claude Sonnet 5

#285

Earlier quoted context omitted.

Okay I don’t care about “eventually”, I want Fable now.

Have you considered getting better at coding so you can build stuff yourself instead of waiting for models you might not be able to get access to anymore?

I'd love to meet the devs who can spin up full feature web apps in under 15 minutes with all the bells and whistles I've gotten Claude to spin up and code. I don't think the AI haters understand the level of time cutting that you can achieve with a very simple and reasonably crafted prompt.

I'm talking back-end, with database models, classes, queries, accompanying front-end layouts, with real dynamic data, running. Stuff that takes days to weeks to spin up, with minimal errors or issues, having cut down on days or weeks of effort, you can focus on testing and making it all into better code.

Re: Claude Sonnet 5

#286

interesting how much worse the sentiment around Anthropic is getting

Seems like a combination of multiple factors: "They took my shit away!" -- 3-day Fable 5 addicts (me) "How dare they tell Trump no?" -- US nationalist / "my country right or wrong" types "Great to see a closed source company fail!" -- open source boosters "Great to see an American company fail!" -- anti-US, and/or pro-China folks "Great to see a successful company fail!" -- anti-capitalists and/or sour-grapes crab bu…

I'm personally in the "they keep releasing shameless lobbying papers disguised as thinly veiled research or essay-coded content, push anticompetitive walled-garden practices, show little else but contempt for their non-enterprise customer base, refuse to communicate about anything and choose public silence as their baseline, seemingly force their employees into vows of public silence as well, actively degrade their products across the board with their vibeslop approach with measurable impacts on customers, openly attack not only open weights models but open source software, and all while pretending they're the 'public benefit corporation' formed by a valiant group of heroes escaping from a duplicitous snake and who, even in light of their own massively duplicitous behavior as of late, should apparently be trusted to be the some sort of arbiter over what this tech should get to be and how it should get to be used while they could hardly be more gleeful about how we're all going to be replaced in 6 months from now perpetually" camp.

Which is a bit of a bummer considering they do genuinely make the best model that's most pleasant to work with in my opinion.

Re: Claude Sonnet 5

#287

Earlier quoted context omitted.

You're the second person that has said this but I cannot understand why you are interpreting the "Agentic computer use" graph in this manner. The graph shows that Opus is cheaper than Sonnet for the same performance. Unless I am suffering a cognitive blindness thing right now.

Wrong! Look at it better. It shows that Opus has superior performance but at higher cost.

Why are you comparing xhigh reasoning between Sonnet and Opus? Of course Sonnet xhigh is cheaper than Opus xhigh, but that isn't the point; the point is that at e.g. 80% accuracy on Opus costs ~$0.45 (medium reasoning) whereas on Sonnet it costs ~$0.52 (xhigh/max reasoning).

Re: Claude Sonnet 5

#288
post #109

Earlier quoted context omitted.

This is why Fable was so good. It followed instructions and it was in no way lazy.

People keep making comments about fable like this? You could only use it for what like a week? How is that at all enough time to evaluate? Opus 4.6 didnt suffer from this problems for a hot minute and then when newer models were released it got worse. I think they change a ton behind the scenes and allocate compute however they want, so the model you use today may behave much differently than how it behaved yesterday

You didn't really have to use it more than a day honestly to tell what kind of shocking paradigm change it was. Man do I miss it.

Re: Claude Sonnet 5

#289

I'm struggling to understand why I'd ever use this instead of just using a lower effort level for opus given on many of the benchmarks listed the cost per task rises above opus at anything higher than medium effort. Only thing I can think of is for when someone is out of opus credits. Of course there are API billing use cases but I'd probably still just use opus on low.

More and more I find myself trying to stop Opus from doing something stupid, and at every turn I need to tell it to stop overcomplicating things. I think the models are being optimized for wealth extraction from users and companies, instead of solving problems. I don't know why Opus would try to create an entire library when I told it specifically to do something simple that would take 2-3 lines of Python.

> I don't know why Opus would try to create an entire library when I told it specifically to do something simple that would take 2-3 lines of Python.

Because it reasons in one direction. First it encounters some kind of issue with 2-3 lines of Python that might make it not work, and then it goes onto plan B, which is making a library, but it doesn't circle back and compare the effort of making the library to working around whatever might make the 2-3 lines not work. Except sometimes it does, because it's inscrutable.

Re: Claude Sonnet 5

#290

Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models. I have been using Sonnet 4.6 more than Opus, because I'm mostly doing agent-assisted development and not fully agent-driven development. This announcement does not make me positive, I have fou…

Yeah, there's a real opportunity for one of these companies to invest time in a model that's tuned for, to use your term, agent-assisted developement. Trouble is, everyone inside their buildings seems to believe that no one will be working like that in a year or two.

They have to, but also everyone working at 3D printing companies thought "industry 4.0" is going to completely override everything, we are going to print housing and going to print a mug at home and drink coffee out of it.

Today's news that Amazon is hiring 11k interns. I think part of the AI story was used as a convenient excuse to get rid of some "fat" and some covid overhiring and gave companies an out to change course.

Post reply on HN