Live data from Hacker News

Apple Silicon costs more than OpenRouter

williamangel.net

201–210 of 322 posts

Re: Apple Silicon costs more than OpenRouter

#201
post #99

Earlier quoted context omitted.

But isn’t training models, a forever task like iterating in tech you can never take a day off, adding humans to the equation don’t humans train/teach themselves new skills over a lifetime, and isn’t one of the selling points in the future when selling this AI slop your AI never goes to sleep and can always be trained forever? The AI price for entry as we go on into the future will only increase.

I agree that training is a forever task, and the current rate of training is probably not sustainable. But all that means is that once the current investment mania ends, the market will most likely find a new equilibrium where continuous training still happens, but at a slower rate that can be sustained by inference revenue.

> but at a slower rate that can be sustained by inference revenue. reply

also it's possible that the scale of inference needed (e.g. Jevons paradox) keeps growing to the point that training costs can fully be absorbed (since training cost is one off vs. inference that can scale).

(I suspect that might be the thinking, I don't know if it will be true, it's also possible that no model will create a moat big enough to attract enough of the inference traffic to make it true).

Depending on the chips/architecture used, the off-peak traffic from inference can also subsidize the training costs.

Re: Apple Silicon costs more than OpenRouter

#202
post #59

Frontier AI companies are selling at a loss. Excusing everything else that u/bastawhiz said[0]; the obvious fact here is that Claude, OpenAI, Gemini et al. are quite literally burning through 100's of billions of dollars and selling it back to you for pennies on the dollar in the hopes that they get to be the only one left. If I spend $10 growing Oranges and sell them to you for $1; then of course it's more expensive…

That seems unlikely. There are many providers for open models on openrouter. It seems unlikely that they are throwing money away for each token they sell. Also, there a good technical reasons for inference being much more efficient at scale.

Sure. And there’s Lyft and Uber and plenty of others. And Grubhub and DoorDash and uber and how many others. And I don’t even fucking remember how many electric fucking scooter companies, I’m practically falling over scooters! I’m sure they aren’t earning market share by selling at a loss either.

Re: Apple Silicon costs more than OpenRouter

#203

Earlier quoted context omitted.

I’ll keep my data local over a $.02/mtok difference.

It’s more than just data locality. OpenRouter is faster, no? I have an M4 pro, and anything but the smallest dumbest models are unusably slow for interactive use. I personally haven’t yet found a good use case for offline/non-interactive LLM work locally.

I played with classifying and summarizing my entire email history (per email) with small models, but that only took about 12h of GPU time at most. Using a coding agent cli wrapper in that case is far slower because of all the spin up cost and the system prompt they inject even if you want to turn it all off.

If I used an actual direct API it probably would've been much faster, but I'm doing it for hobby / fun reasons. You also get to fiddle with a lot more params.

Re: Apple Silicon costs more than OpenRouter

#204
post #112

If you want a good dense model, use qwen3.6 27B instead, speed will be up, and if you don't take my word for it being smarter, take openrouter's prices of it against the bigger, slower and less memory-efficient gemma do the talking. If you want a faster model, go for qwen3.6 35B (or gemma 4 26B if gemma models perform better for your tasks). There is a reason why people (myself included) haven't shut up about those t…

Not disagreeing with your argument, but: > If you want a good dense model, use qwen3.6 27B instead, speed will be up, and if you don't take my word for it being smarter, take openrouter's prices of it against the bigger, slower and less memory-efficient gemma do the talking. Don't know if this is the correct read. I think those providers are simply taking cue from Alibaba's first-party pricing for the 27B Dense. It's…

I feel like if I had the infrastructure and saw that there is a huge interest in the model, i'd just undercut alibaba's prices a little harder to grab all the consumers. I am sure that the providers have done the math and found that there is a reason not to do this (compute-bound if too many users?), but the delta is very stark, especially for output. Last I checked the cheapest 27b on openrouter was 2$ out vs 0.38$ for the 31b.

But I do agree that the openrouter prices aren't a strong signal and probably should have worded it a little better. It's just a really stark and 'in your eyes' gap.

Re: Apple Silicon costs more than OpenRouter

#205

This isn't a good analysis, and it's because it keeps rounding everything up. He rounds up the cost of electricity by 10%. He has a range of power use, takes the high end (which is 2x the low end) and multiplies it by the inflated electricity cost. But then they talk about using a newly purchased Mac to do the inference, running at full capacity, 24/7. Why would you do that? Apple silicon is fast but the author point…

Rounded up, yes, and oddly inefficient for someone obsessed with inefficiency. One could buy a brand new 64gb M5 macbook for well over 4k. Another could buy a scratched up but functioning M1 Max 64gb off of ebay for a little over 1k—and somehow get the same 10-20 t/s with 31b that the author does with an M5. Or better yet, have a frontier model do the planning and judging, and have a local MOE model execute at 50 t/s…

I have an M1 Pro, and a M4 & M5 max to play with at work and the speed difference is very significant between all 3 machines, the M1 Pro is far slower, and the M5 is significantly faster than the M4. And a windows 3090 beats all of them but eats twice the amount of power per token. This is all running the same 24GB memory friendly model with LM studio.

Re: Apple Silicon costs more than OpenRouter

#206

This isn't a good analysis, and it's because it keeps rounding everything up. He rounds up the cost of electricity by 10%. He has a range of power use, takes the high end (which is 2x the low end) and multiplies it by the inflated electricity cost. But then they talk about using a newly purchased Mac to do the inference, running at full capacity, 24/7. Why would you do that? Apple silicon is fast but the author point…

Rounding everything down in the most optimistic setting got me to $0.40 per million tokens, and openrouter has the same model at $.38/mtok.

Also many have power even cheaper or even free unused surplus power with solar.

I don't do local inference other than hobby & learning reasons because electricity is so expensive where I am at.

Re: Apple Silicon costs more than OpenRouter

#208
post #71

Earlier quoted context omitted.

Profitable for inference if you completely ignore training costs and that you absolutely must continuously train new models.

Which is where your analogy breaks down and why you think you’re taking crazy pills. Inference is growing and selling the oranges in your analogy. Model building is growing the farm to sell larger, juicier more addicting oranges.

Last month Anthropic tried to control the narrative by drumming up the “super scary AI” trope.

The news they successfully buried was that companies like AirBnB are now running Qwen and open source models. The free oranges are now good enough. There is no future unless the goal is to get to super intelligence and utterly take over the world before anyone else gets one. Anything else and free models are six months behind. The money now is the opposite of what everyone thought a year ago: datacenters. Everyone thought AWS was fucked. Turns out AWS is really good at running Qwen.

Re: Apple Silicon costs more than OpenRouter

#209
I suppose folks here already know this but it deserves a mention: subscription pricing is 10-20x cheaper than API pricing at Anthropic for example and it will be a far better experience (better models, faster responses, as much parallelism as you want, etc) so if it works for you there's no economic argument to buy a machine for inference at the moment.

Re: Apple Silicon costs more than OpenRouter

#210
post #176

Earlier quoted context omitted.

Articles like that still miss a bit of the nuance. Imagine having your house paid for, and you grow old and you have no rent to pay. Yes, you could have invested but likely you would have spent some of that money on something else, or your investments might have not worked out so well, or any other reason. Human reasons, to be specific. Owning property is like a lock.

Imagine having your house paid for, and you grow old and you have no rent to pay. My home is "paid for". Except for the HOA and property taxes that are not that far off from what I was previously paying in rent, the ongoing maintenance costs with random large spikes, and the opportunity cost of having a large chunk of money in the house and not in the market. It was still probably the right decision, but it's not at…

Surely though, the HOA and all that would likely be baked into a renter's price.

And you didn't need to go live in a HOA. I don't, and it's much cheaper.

Post reply on HN