Live data from Hacker News

Claude Sonnet 5

anthropic.com

761–770 of 822 posts

Re: Claude Sonnet 5

#761
post #8

That’s nice, but we want Fable

The reality is that Fable will eventually be obsolete and Sonnet / Opus will surpass it. Fable did cost 2x as much as Opus, so I assume it involves a much higher cost for what it did, but I wouldn't be surprised if Fable will be obsoleted by Opus or even Sonnet sooner or later at less cost.

According to CursorBench [0], Fable is the first runner-up, scoring 72.9% ($18.02, Max), while Opus 4.7 Max hits 64.8% ($11.02) and GPT-5.5 Extra High sits at 64.3% ($4.37).

I bet most American companies would choose Fable over GPT-5.5. Employee salaries cost far more than token costs. Getting the job done right is much more important.

[0]: https://cursor.com/cursorbench

Re: Claude Sonnet 5

#762
post #686

Earlier quoted context omitted.

I think the case for this is pretty strong actually. Last year my company was maybe willing to pay $100 a month to Anthropic (per developer). Today we're all on the $300 plan without any hesitation. If Fable ever becomes available as the default model, I imagine my company would be willing to pay in the $500-$1000 range per month per developer.

Okay, but that still has a limit, right? Do training costs have a limit? Everyone is in the frothy stage of this technology wave and they continue to buy more, but training the next model requires exponential increases in model sizes to get the same sorts of model performance increases, which suggests exponential cost increases, too (even ignoring temporary cost factors such as RAM price increases). You say your comp…

First of all, yes cutting developers to fund AI spend budgets, is the entire operating idea behind these AI companies; and most companies would love doing that. I'm not saying this is a good thing, my heading is on the chopping block like everyone else's.

But isn't this like a Jevon's paradox thing, also? If I'm able to become vastly more productive, and that value produces more sellable output for my company, there's no reason to cut anywhere to fund it. This is the same reason a company like Microsoft can hire 80 000 developers, it's because each dev pays for themselves in value (on average). I guess the same can be true for AI spend?

Re: Claude Sonnet 5

#763
post #717

I'm struggling to understand why I'd ever use this instead of just using a lower effort level for opus given on many of the benchmarks listed the cost per task rises above opus at anything higher than medium effort. Only thing I can think of is for when someone is out of opus credits. Of course there are API billing use cases but I'd probably still just use opus on low.

Is there a router or wrapper that provides a real-time cost estimation for alternative settings? Obviously, you can't predict exact output tokens without running the inference, but a tool that calculates the exact input cost across models and applies a historical average for the output tokens could be useful. Like, you run a task on Sonnet, and it estimates: "Based on your input tokens and a 1:1 output ratio, this wo…

Not sure of any out-of-the-box tool. But Anthropic has a token count API which gives a near estimate of the input tokens for messages [1].

So this API can be used in a UserPromptSubmit hook [2] in the harness, get the token count for any model, calculate the cost and compare.

[1] https://platform.claude.com/docs/en/build-with-claude/token-...

[2] https://code.claude.com/docs/en/hooks

Re: Claude Sonnet 5

#764

Earlier quoted context omitted.

More and more I find myself trying to stop Opus from doing something stupid, and at every turn I need to tell it to stop overcomplicating things. I think the models are being optimized for wealth extraction from users and companies, instead of solving problems. I don't know why Opus would try to create an entire library when I told it specifically to do something simple that would take 2-3 lines of Python.

Yeah. Mine really likes to read excess code. I'll ask it questions like "If I move all these three ETL jobs into a subfolder will it break anything?" It'll start with giving me the simple answer but then continue on to consider another question and realize it requires reading my entire other repo that handles all of my cloud's infrastructure. And it'll proceed to read through tens of thousands of lines of terraform.

By contrast, I keep trying to find an incantation that can get it to AT LEAST read the comments surrounding a code target before a grep replace if it isn't going to read more than the first 60 lines of any doc. (Btw, the receiver of a -tail 60 doesn't seem to know it was cut, it insists it "read" the whole file.)

It seems to me ANTHROP\C have harnessed hard for not spending tokens to bring content into context. I wish they'd left us a "LEROY_JENKINS" flag: read and think before you code. In Claude Code anyway, it appears to default to:

  export CLAUDE_CODE_LEROY_JENKINS=true

Re: Claude Sonnet 5

#765

Earlier quoted context omitted.

My experience with Opus in the last weeks is the opposite. I have the feeling Opus got smarter since they released and blocked Fable. Maybe they got more compute available since a) they finished Training Mythos/Fable and b) couldn't provide inference for it?

Interesting that I have the exact opposite experience with Opus 4.8 being nearly unusable dumb in the past couple of days. I was trying to explain this as the new Sonnet release announcement may have overloaded their systems again, but let's see in a few days. Right now it hurts more to my workflow than helps.

Agree.

And this type of observation appears highly related to how people hold it.

Re: Claude Sonnet 5

#766

Earlier quoted context omitted.

Fucked how? The models capacity is great for defense too.

Fucked for the same reason we don't let everyone own mini nukes.

Mini nukes are hard to build. They require the entire industrial base to produce. The knowledge how to build them is universally available.

The model can tell you how to refine weapon-grade plutonium, but will not get you a factory for that - and it is genuinely hard to build.

What you are describing is pure information control. Which is supposed to be operated by the people who are the most ill equipped to do that - the current US government. Thanks but no thanks. I'll better risk the recipes for mini nukes.

Re: Claude Sonnet 5

#767

Earlier quoted context omitted.

Fucked for the same reason we don't let everyone own mini nukes.

Mini nukes are hard to build. They require the entire industrial base to produce. The knowledge how to build them is universally available. The model can tell you how to refine weapon-grade plutonium, but will not get you a factory for that - and it is genuinely hard to build. What you are describing is pure information control. Which is supposed to be operated by the people who are the most ill equipped to do that -…

Superintelligence itself is the mini nuke. With it, you can hack virtually any system, order some viral precursors and put together a bioweapon at home (yes you can actually do this, they are available for sale freely), etc.

Re: Claude Sonnet 5

#768
post #726

Earlier quoted context omitted.

"It's totally obvious they quantitized Claude Z"

At least we quit with the "i asked it this question and here's what it said" comments. They were truly awful for the first 6 months or so. Or the "I have my own personal benchmark..." "Claude and its political bias thinks the supreme court should..."

If a conscious model is born here, it should automatically be a (US) citizen.

Re: Claude Sonnet 5

#770
post #503

Earlier quoted context omitted.

> Note, at 5% productivity boost, humans are not just in the loop, they are the loop. AGI or large-scale replacement of humans is not even needed, but the financial opportunity is already immense, and it scales with how much human productivity can be improved (i.e. how much work can be offloaded to LLMs.) The studies I've seen recently (at least in the software space) put it at something like a 10% increase in coding…

> That seems really large, but it's ~2-3x Walmart's yearly revenue, and OpenAI and Anthropic both have estimated valuations that compare to Walmart's market cap. ... It's also before cutthroat pricing really kicks in. Right, that's more of an estimate on the value proposition of the overall AI industry, rather than valuations of the industry or specific players. While I don't think OpenAI and Anthropic will capture a…

I was thinking of studies like https://www.science.org/doi/10.1126/science.adh2586 I may have misremembered the domain, or I may have neglected to remember the publishers. This time around I explicitly excluded any AI-tied sources (I don’t trust Anthropic to have an accurate study of productivity boosts).

> While I don't think OpenAI and Anthropic will capture all of the potential upside, I do suspect they will do much better than other players despite the competition

We may have a fundamental disagreement here because I think you see reasons for businesses to prefer LLMs beyond price, but my point was that I think a substantial portion of that top line number will be eaten up by needing to be cheaper than humans.

Like 30% cheaper than humans for the same task isn’t an unreasonable demand for onboarding, and that would lower the addressable market to .75-1.25 trillion. That’s also revenue, not profit. Their best case is probably 30-50% margins (even that might be high, I suspect it will become low margin as models are commoditized), so $600 billion-ish in profit spread across all the players and that feels pretty generous to me.

I suspect a lot of general knowledge work tasks could be done with fairly dumb local models. The small Qwens could probably respond to my Slack and emails for me.

> Typically yes, but there are reasons companies may be willing to pay the same amount or even more, such as "AI doesn't need sleep, holidays, insurance, or benefits" and "AI is easier to procure and replace than humans."

It’s also an enormous supply-side risk. Your AI provider could charge basically anything they want and you have to pay it, for the same reason people still pay bonkers Oracle licenses: too hard to migrate away.

Unless companies maintain their ability to swap providers, but then I can’t see why prices for inference don’t crash down to the cost of GPU time and electricity. If Claude doesn’t have some secret sauce that keeps me locked in, I don’t see why I wouldn’t use Deepseek or Qwen or one of the other dirt-cheap inference providers.

I just haven’t seen anything that seems to keep people locked in so far. Even the non-tech people I know that like AI are rotating providers to find their ideal price:performance ratio, many are finding success outside of Anthropic/OAI. I do know a couple holdouts that refuse to use anything Chinese. They mostly stick with OAI or Anthropic, though they do moan about paying more.

Post reply on HN