Earlier quoted context omitted.
> You think AWS are going to subsidise your usage of somebody else’s models indefinitely? As with Costco's giant $5 roasted chickens, this is not solid evidence they're profitable. Loss-leaders exist.
Rather than speculating another option is to just measure things. I churned through billions of tokens for evals and synthetic data earlier this year, so I did some of that. On an H100 node, a Llama3 70B FP8 at concurrency=128 generated at about 0.4 J/token (this was estimating node power consumption and multiplying by a generous PUE, 1.2X or something like that) - it was still 120X cheaper than the 48 J/token estima…
LLMs are cheap
151–160 of 319 posts
Re: LLMs are cheap
#152Earlier quoted context omitted.
I’m talking about today’s experience, not speculating about what might happen at some arbitrary point in the future. The web has this shitty UX. LLMs do not have this shitty UX. I’m going to judge on what I can see and use.
> I’m talking about today’s experience… In that case, get uBlock. The answer is in the first result, on the first screen, and the answer is even quoted in the short description from the site. (As a bonus, it also blocks the cookie consent popups on the AA site, if you like.) The only thing getting in the way of the real, vetted, straight-from-the-source answer currently is the AI overview . https://imgur.com/a/pRUGgR…
Even so, saying that the UX of the web is almost as good as the UX of an LLM after you take steps to work around the UX problems with the web isn’t really an argument.
Re: LLMs are cheap
#153Earlier quoted context omitted.
> I’m talking about today’s experience… In that case, get uBlock. The answer is in the first result, on the first screen, and the answer is even quoted in the short description from the site. (As a bonus, it also blocks the cookie consent popups on the AA site, if you like.) The only thing getting in the way of the real, vetted, straight-from-the-source answer currently is the AI overview . https://imgur.com/a/pRUGgR…
Most people don’t use an ad blocker. Even so, saying that the UX of the web is almost as good as the UX of an LLM after you take steps to work around the UX problems with the web isn’t really an argument.
I mean, they should. Anyone on this site most certainly should.
The LLM UX is going to rapidly converge with the search UX as soon as these companies run out of investor funds to burn. It's already starting; https://www.axios.com/2024/12/03/openai-ads-chatgpt.
What then?
Re: LLMs are cheap
#154You can't compare an API that is profitable (search) to an API that is likely a loss-leader to grab market share (hosted LLM cloud models). Sure there might not be any analysis that proves that they subsidized, but you also don't have any evidence that they are profitable. All the data points we have today show that companies are spending an insane amount of capex on gaining AI dominance without the revenue to achiev…
> an API that is likely a loss-leader to grab market share (hosted LLM cloud models) Everyone just repeats this but I never buy it. There is literally a service that allows you to switch models and service providers seamlessly (openrouter). There is just no lock-in. It doesn't make any financial sense to "grab market share". If you sell something with UI, like ChatGPT (the web interface) or Cursor, sure. But selling…
Re: LLMs are cheap
#155And now?
Re: LLMs are cheap
#156Earlier quoted context omitted.
This is addressed in the article. Giving arguments for llms being profitable as APIs.
It's addressed poorly. > First, there's not that much motive to gain API market share with unsustainably cheap prices. Any gains would be temporary, since there's no long-term lock-in, What? If someone builds something on top of your API, they're tying themselves to it, and you can slowly raise prices while keeping each increase well below the switching cost. > Second, some of those models have been released with ope…
That's not really how the LLM API market works. The interfaces themselves are pretty trivial and have no real lock-in value, and there's plenty of adapters around anyway. (Often first-party, e.g. both Anthropic and Google provide OpenAI-compatible APIs). There might initially have been theories that you could not easily move to a different model, creating lock-in, but in practice LLMs are so flexible and forgiving about the inputs that a different model can be just dropped in an work without any model-specific changes.
> 80% margin on GPU cost? What about after paying for power, facilities
The market price of renting that compute on the market. That's fully loaded, so would include a) pro-rated recouping the capital cost of the GPUs, b) the power, cooling, datacenter buildings, etc, c) the hosting provider's margin.
> admin, support, marketing, etc.? Are GPUs really more than half the cost of this business?
Pretty likely! In OpenAI's leaked 2024 financial plan the compute costs were like 75% of their projected costs.
Re: LLMs are cheap
#157You can't compare an API that is profitable (search) to an API that is likely a loss-leader to grab market share (hosted LLM cloud models). Sure there might not be any analysis that proves that they subsidized, but you also don't have any evidence that they are profitable. All the data points we have today show that companies are spending an insane amount of capex on gaining AI dominance without the revenue to achiev…
There is also a lot of different models at a lot of different price points (and LLMs are fairly hard to compare to begin with). In this theory of a likely loss-leader, must we assume that all of them, from all companies, are priced below cost...? If so, that seems like a fairly wild claim. What's Step 2 for all of these companies to get ahead of this, given how model development currently works? I think the far more…
Re: LLMs are cheap
#158Earlier quoted context omitted.
Rather than speculating another option is to just measure things. I churned through billions of tokens for evals and synthetic data earlier this year, so I did some of that. On an H100 node, a Llama3 70B FP8 at concurrency=128 generated at about 0.4 J/token (this was estimating node power consumption and multiplying by a generous PUE, 1.2X or something like that) - it was still 120X cheaper than the 48 J/token estima…
Yes, if we discount the billion or so Facebook spent to train Llama3.
If you would like to someone add that somehow as a line item, perhaps you should add the full embodied energy cost of Linux (please include the entire history of compute since it wouldn't exist without UNIX), or perhaps the full military industrial complex costs from the invention of the transistor? We could go further.
Re: LLMs are cheap
#159Earlier quoted context omitted.
The entire businessmodel may only work as long as inference takes up the physical space and cost of a small building. Last time personal computing took up an entire building, we put the same compute power into a (portable) "personal computer" a few decades later. Can't wait to send all my data and life to my own lil inference box, instead of big tech (and NSA etc).
“Last time” we weren’t up against physical limitations for solid state electronics like the size of an atom, wavelength of light, quantum effects, thermal management, etc.
while few years back, it do it bianually
Re: LLMs are cheap
#160For eg. Claude was undoubtedly the best model for software devs until gemini 2.5 was released and now i see people divided with majority of them leaning towards Gemini.
And there is very little room for mistakes, as we have seen how llama became completely irrelevant in matter of months.
So while inference in itself can be profitable (again thats a big *), these companies will have to keep fighting for what it looks like decades unless one of them actually solves hallucinations and re constructs computer interfacing at a global scale!