Live data from Hacker News

Apple Silicon costs more than OpenRouter

williamangel.net

251–260 of 322 posts

Re: Apple Silicon costs more than OpenRouter

#251
post #243

Earlier quoted context omitted.

Your post makes sense if you bought the hardware for other reasons, and maybe run models occasionally as a novelty. That isn't the case for many, though, and there is a whole social media space where people are hyping up the latest homebrew options for running models, believing it frees them from the yoke of big AI. Millions of people are buying big $ maxed-out hardware like the Mac Studios or DGX specifically to run…

> Millions of people are buying big $ maxed-out hardware like the Mac Studios or DGX specifically to run LLMs. What's your source for this?

Not just one, but two replies completely hung up on that. Like, why even reply when you already saw the other guy doing the same tired pedantry? Just wanted to feel like you contributed?

Ignoring that it was just tosser hyperbole (that absolutely zero reasonable people need to question), yes, enormous numbers of people are buying GPUs or hardware with the explicit goal of running local LLMs, and social media is full of people hyping various setups and models. Mac Minis are almost impossible to find, and that alone is selling at a clip of about 300,000 every four months. Large memory GPUs are basically a myth at this point. All so people can pay more to get a worse result than commercial options, which is precisely the point of the submission.

These local setups only ever make sense if you have something that confidential, or you're doing something that ToS of the majors would ban you for.

Now given this pedantry horseshit, you'll probably demand that I specifically show a citation on DGX or Studio sales, which...rofl.

Re: Apple Silicon costs more than OpenRouter

#252

This isn't a good analysis, and it's because it keeps rounding everything up. He rounds up the cost of electricity by 10%. He has a range of power use, takes the high end (which is 2x the low end) and multiplies it by the inflated electricity cost. But then they talk about using a newly purchased Mac to do the inference, running at full capacity, 24/7. Why would you do that? Apple silicon is fast but the author point…

The real reason this comparison makes no sense is that only a vanishingly small fraction of people seriously using ai to code would seriously use a model so far from the top models (including open source ones).

He should compare his MacBook to Open Router on Kimi 2.6 1.1T or GLM 5.1 (754B), at bfloat16 precision, which he can't ofc.

But it furthers his point that things like open router are a better idea, which is not surprising.

Re: Apple Silicon costs more than OpenRouter

#253
post #82

Earlier quoted context omitted.

Which is where your analogy breaks down and why you think you’re taking crazy pills. Inference is growing and selling the oranges in your analogy. Model building is growing the farm to sell larger, juicier more addicting oranges.

Are ya fuckin' serious mate? The restaurant next to the mines were profitable up until the moment the mines themselves shut down: one doesn't exist without the other. You can't ringfence inference as "the profitable bit" and then hand-wave away the training. Without continuous training there is no inference product. Claude 3 Opus isn't sitting there making revenue in 2026 - the thing is just deprecated . The moment y…

Dario has stated they made more from selling inference to Opus 3 than they did training it. Same with 3.5, 4.0 and I'm assuming 4.5 as well.

The reason they're losing money on paper is because the models keep growing 10x in size every generation but they're not getting 10x returns on model inference (closer to 2x)

Re: Apple Silicon costs more than OpenRouter

#254
post #82

Earlier quoted context omitted.

Are ya fuckin' serious mate? The restaurant next to the mines were profitable up until the moment the mines themselves shut down: one doesn't exist without the other. You can't ringfence inference as "the profitable bit" and then hand-wave away the training. Without continuous training there is no inference product. Claude 3 Opus isn't sitting there making revenue in 2026 - the thing is just deprecated . The moment y…

Dario has stated they made more from selling inference to Opus 3 than they did training it. Same with 3.5, 4.0 and I'm assuming 4.5 as well. The reason they're losing money on paper is because the models keep growing 10x in size every generation but they're not getting 10x returns on model inference (closer to 2x)

I know what you’re quoting, but you've badly misread it.

Dario's point was the opposite of yours; he used per-model accounting to explain why the company P&L gets worse every year, not better. His own numbers (10x training costs each generation and ~2x revenue return.)…

"It looks like it's getting worse and worse" are his words, not mine.

In 2025, Anthropic's inference costs came in 23% over their own projections. They cut their gross margin forecast from 50% to 40%.

Re: Apple Silicon costs more than OpenRouter

#255

Earlier quoted context omitted.

Let's not get ahead of ourselves. Millions , really? I can believe there are a lot of enthusiasts doing this, but "millions" needs a citation.

This is HN; it has probably never occurred to half the people here that the average person even in first world countries doesn't even have the financial capacity to make an impulse five-figure USD purchase, even if on credit.

No one said anything about this being an "impulse" purchase -- it would usually be perceived as a "career investment" purchase, given that many feel they need to race to be with "it" or be left behind -- nor does "five figure" have anything to do with this -- DGX options are available at $3000 -- there's a certain irony that you are posting a comment that is basically "people in my country/circle can't, therefore no one can", while dismissing my comment that many people can.

Millions of people are paying thousands of dollars a year to buy a slightly upgraded entertainment package in their car. There are 60 million or so millionaires alone, including 6 million+ in China.

There are a lot of people with a lot of wealth on the planet. A lot. Millions...it isn't that unfounded, friend.

So doing this "this is HN" snide jerk act, and then basically projecting your lot on the planet is...I don't know if you intended it, but it's rather amazing.

Re: Apple Silicon costs more than OpenRouter

#256

This isn't a good analysis, and it's because it keeps rounding everything up. He rounds up the cost of electricity by 10%. He has a range of power use, takes the high end (which is 2x the low end) and multiplies it by the inflated electricity cost. But then they talk about using a newly purchased Mac to do the inference, running at full capacity, 24/7. Why would you do that? Apple silicon is fast but the author point…

Not sure where 40 tokens per second is coming from. I’ve seen 95-100 tokens per second on M5 Max 128GB running Gemma 4 31B. I’ve done experiments where it is faster than Claude Opus 4.5 for the same prompts.

Re: Apple Silicon costs more than OpenRouter

#257

Earlier quoted context omitted.

That seems unlikely. There are many providers for open models on openrouter. It seems unlikely that they are throwing money away for each token they sell. Also, there a good technical reasons for inference being much more efficient at scale.

Sure. And there’s Lyft and Uber and plenty of others. And Grubhub and DoorDash and uber and how many others. And I don’t even fucking remember how many electric fucking scooter companies, I’m practically falling over scooters! I’m sure they aren’t earning market share by selling at a loss either.

Uber has been profitable since 2023.

Re: Apple Silicon costs more than OpenRouter

#258
post #249

Earlier quoted context omitted.

Right. Qwen 3.6 45b (6 parameter) runs on a commodity 5090, which, if you're into video games, you probably already have one of. It is entirely usable for most code generation tasks. (Not all, but most.) Likewise, DeepSeek V4 Flash is quite accessible on local models, with DwarfStar 4 making it easy to run on a 96GB MacBook. There's nothing wrong with paying for inference, but local models bring up some pretty amazin…

Ah yes, most people who play any sort of video game own a $2000 graphics card :p

[dead]

Re: Apple Silicon costs more than OpenRouter

#259

This isn't a good analysis, and it's because it keeps rounding everything up. He rounds up the cost of electricity by 10%. He has a range of power use, takes the high end (which is 2x the low end) and multiplies it by the inflated electricity cost. But then they talk about using a newly purchased Mac to do the inference, running at full capacity, 24/7. Why would you do that? Apple silicon is fast but the author point…

Actually, figuring it on generating tokens 24/7 is the best case scenario. if you figure it at 8 hours a day of actual use, you still have the fixed cost of the hardware being the highest portion of the budget, but now you generate 1/3 the tokens so you triple that cost per token.
Post reply on HN