Live data from Hacker News

Claude Sonnet 5

anthropic.com

371–380 of 822 posts

Re: Claude Sonnet 5

#371
post #253

Earlier quoted context omitted.

Some napkin math -- total global labor compensation is about 50% of the GDP, which puts it in the USD 50 - 60 Trillion range: https://ourworldindata.org/grapher/labor-share-of-gdp This source claims that knowledge workers alone (probably because they are paid much more) account for 35 - 50 Trillion of that: https://github.com/danielmiessler/Substrate/blob/main/Data/K... If LLMs can boost their productivity even by an…

I am deeply surprised by the silence of philosophers, sociologists, liberal arts majors, economists. Where are the think tanks who contemplate and debate the societal aspects? The tech is advancing full steam but the "other side" doesn't feel anywhere nearly ready.

Silence? Even the pope has come out against AI? Who hasn't? Diplo??

Re: Claude Sonnet 5

#372

The cost per task chart is telling me that I should _never_ use Sonnet 5 above medium effort level - Opus always performs better for a given cost. So I guess the takeaway is that if Sonnet 5 medium isn't good enough for you, switch models, not effort levels.

While I appreciate, they publish this information, it's increasingly hard to keep track of it all. I've lost the mental model of how different models at different effort levels perform and what tasks they are good at. In practice, I tend to just use the default on Claude Code that works well enough. But I wonder to what degree other users really play around with these settings to optimize for their project.

It's almost like you want an automatically intelligent choice of your artificial intelligence.

Understandable frankly.

Re: Claude Sonnet 5

#373
post #313

Earlier quoted context omitted.

Idk why you're perceiving silence. Feels to me like this is the main thing people talk about nowadays.

It has to do with the scope of what they're discussing. It seems extraordinarily small: e.g. what if AI increases productivity growth by 0.4%? Do data centers use too much water? Are AIs racist when reviewing resumes? The frontier labs, on the other hand, are thinking about replacing all human labor, ending death, and the risk of it causing human extinction. Most of the apparatus we're talking about approach it very…

The public would happily string up any of these CEOs if given the chance

Re: Claude Sonnet 5

#374
post #246

Earlier quoted context omitted.

The insiders disagree because they are benefiting greatly from the insane valuations, right? Chinese models are quickly commodifying frontier inference, the US Gov is preventing domestic SOTA models access to the public and without those models why would consumers still spend $200/month to use the best models? It’s such a mess and isn’t inspiring confidence as a non-investor.

Are they benefiting from the insane valuations though? If the valuations deflate before the insiders are able to exit, I think that would be worse for them than a lower but sustainable valuation. It all comes down to whose prediction of the future is closer to correct. I think the most likely future is commodification of inference and "agent-assisted" rather than "agent-driven" workflows dominating the future of work…

Even if the future is agent-driven workflow, that doesn't stop the commodification of inference. a good agent-driven workflow, in my experience, is a byproduct of the harness and scaffolding around the agent.

What insiders are you talking about? They're going to be hot towards the possibilities so they can exit to a massive windfall. I dont know why they would want to be publicly critical of these technologies that could make millions on IPO.

Re: Claude Sonnet 5

#375

Earlier quoted context omitted.

It seems to be more them losing goodwill combined with their marketing. I don't agree with your framing that all negativity is from crazies

I don't think all the negativity is from crazies, but big chunks of it are certainly motivated. I certainly left out numerous other categories.

The amount of anti-Anthropic and anti-Dario posts i've seen on reddit threads has gotten a bit ridiculous.

It feels like your analysis is mostly spot on, it's the confluence of several motivated parties pouring effort into social media.

Many of the posters are pro-foreign models/pro-open source, and most can't distinguish the difference between "open source" and open weight models like Qwen, Minimax, or GLM.

Reminds me of the old "free as in beer" vs "free as in speech" debate. Free beer means you don't pay, but you don't get to see the recipe or change it. Free speech means you get the actual source and the right to study it, modify it, and redistribute it.

Open weight models are basically the beer version. You can download the weights, run them locally, fine-tune them, quantize them, host them on your own boxes — but what you have is a finished product, not the blueprint for how it was built.

Re: Claude Sonnet 5

#376
post #109

Earlier quoted context omitted.

This is why Fable was so good. It followed instructions and it was in no way lazy.

People keep making comments about fable like this? You could only use it for what like a week? How is that at all enough time to evaluate? Opus 4.6 didnt suffer from this problems for a hot minute and then when newer models were released it got worse. I think they change a ton behind the scenes and allocate compute however they want, so the model you use today may behave much differently than how it behaved yesterday

For me claude-fable-5 failed to follow the instruction following test I'm making against various models https://github.com/marcindulak/claude-fails-to-follow-claude...

Re: Claude Sonnet 5

#377
post #315

Earlier quoted context omitted.

While I appreciate, they publish this information, it's increasingly hard to keep track of it all. I've lost the mental model of how different models at different effort levels perform and what tasks they are good at. In practice, I tend to just use the default on Claude Code that works well enough. But I wonder to what degree other users really play around with these settings to optimize for their project.

Exactly this is my problem with all AI tools. I want someone else to create working tools for me so I can focus on my product. It is the same with other tools. I do not want to spent huge amounts of energy and time to setup my IDE, operating system or desk layout. I guess it is too early to have that now.

I think that's the whole selling point of lovable?

Re: Claude Sonnet 5

#379
post #227

I'm struggling to understand why I'd ever use this instead of just using a lower effort level for opus given on many of the benchmarks listed the cost per task rises above opus at anything higher than medium effort. Only thing I can think of is for when someone is out of opus credits. Of course there are API billing use cases but I'd probably still just use opus on low.

Older Opus models will likely get deprecated and then over time this is the cheapest model. That is how prices are currently increased.

Yeah... Sonnet becomes the new cheap model, and some Fable class model becomes the more expensive/better one.

Re: Claude Sonnet 5

#380

Earlier quoted context omitted.

What I want is a harness that knows how to optimize this kind of thing for me.

Which is your own harness and your own evals for your tasks I guess

I don't demand a customized compiler for my code even if such a compiler could outperform gcc. There is a lot of value in focusing on correctness to an extreme degree even if the outcome might be suboptimal to something more tailored - a tool with a large customer base can justify more resources going into its maintenance.
Post reply on HN