Live data from Hacker News

The real prices of frontier models

playcode.io

31–40 of 91 posts

Re: The real prices of frontier models

#32
It's a bit unrelated, but I've been wondering if LLM providers are using cache read costs to preserve the illusion of constant output token prices. In reality, the longer your context the more expensive output tokens are, but anthropic and openai both have flat output token pricing.

In practice though as a result of cache reads over multiple turns you will end up paying quadratic pricing anyway.

Re: The real prices of frontier models

#33
post #17

An individual token, and the level of energy it represents (electricity, or relative effectiveness per model) increasingly seems the space of obsfucation. This space can be increasingly avoided by becoming, and remaining, efficient and effective with prompts.

That is one of the interesting things about Neuralwatt cloud. Their pricing is based on energy rather than tokens (actually they have a token-based alternative, but claim the energy pricing results in 95% cheaper results). I've tried out their subscription offer and it does seem like you get a lot more usage even on the cheap plan. However since the energy metering is pretty much unique to them (at least that I've seen), there really isn't anything to compare it against, so hard to tell how accurate it is.

Regardless, it is cool to be able to contextualize the actual spend in terms of physical energy utilization. It even has a little co2 number (though again, kind of a "trust me bro" metric).

Re: The real prices of frontier models

#34

I have reduced usage of Fable and Sonnet 5 to a minimum. Fable in particular is amazing at creative tasks, but not worth the cost for almost everything else. I can have Opus 4.6/4.7 running non-stop without hitting quota, vs maybe 20 minutes of Fable usage.

After getting used to Opus 4.8, Sonnet 5 is nigh unusable for coding. I much prefer Opus 4.8 + DS V4F for routing. Sonnet 5 is just not useful in the price/performance anywhere. I still get some use of Fable for planning because it can comprehend agent-built codebases (which have repetition of local patterns to a grand degree).

Re: The real prices of frontier models

#37
post #3

The fact that OpenAI documents theirs is already a big improvement over Anthropic. But, also, the OpenAI tokenizer got more efficient when they last updated it, rather than less. https://mdstudio.app/o200k-base-tokenizer

Interesting. New models are estimated at ~5T params, so 45,000x increase over BERT base (110m). But vocab size of 200k, so only an increase of 7x over BERT base (30k).

Where do these estimates come from?

Re: The real prices of frontier models

#38

I have reduced usage of Fable and Sonnet 5 to a minimum. Fable in particular is amazing at creative tasks, but not worth the cost for almost everything else. I can have Opus 4.6/4.7 running non-stop without hitting quota, vs maybe 20 minutes of Fable usage.

Well, I both agree and disagree with you. On one hand, the price is just astronomical for Fable, well, not exactly astronomical, but I would say unaffordable. That is to say, so expensive that it is impossible to use. But on the other hand, Fable is simply incomparable to anything else. I mean, it is just amazing. There is nothing even close to being equal to it.

I have to agree on fable being in another league. I’ve given it my most challenging problems and almost always comes back with a functional, solution 2-5 prompts away from a finished pr. Literally smashed our backlog - very impressed. What i found most efficient is to add “use sonnet agents for research” gets you really far, and on large not so novel tasks “use opus for tasks” by adding this it spins up many agents, works for 2+ hours in a usage window and completes A LOT of work.

Re: The real prices of frontier models

#39
Is it on topic to complain about the various claude-isms in this article? I don't know any actual humans that write titles like "Two floors the rate card hides".

I find my brain disengages once I suspect something of being written by an LLM. If the author didn't put much effort into writing it, should I expect them to have put much effort into fact-checking it?

Edit: this specific title has been deleted from the article. That was not my point! Please put in more effort into writing things that you want others to read! Rather than putting in low effort but being better at hiding it.

Re: The real prices of frontier models

#40

Earlier quoted context omitted.

Interesting. New models are estimated at ~5T params, so 45,000x increase over BERT base (110m). But vocab size of 200k, so only an increase of 7x over BERT base (30k).

Where do these estimates come from?

a dark anal cavity somewhere next to marketing.
Post reply on HN