In my testing, qwen-3.6-27b in full precision is well below sonnet, but above claude haiku in coding tasks. Gemma is not even close to qwen, it’s much, much worse.
Apple Silicon costs more than OpenRouter
271–280 of 322 posts
Re: Apple Silicon costs more than OpenRouter
#272I expect self-hosted to be quite competitive pretty soon. Github Copilot is already wildly more expensive than it was last month. People are going from spending a few bucks to a few thousand for that same usage. So, if it doesn't get a lot more efficient (like 3x the tokens, or more, from the same infrastructure), the prices will have to go up quite a lot to keep the lights on. Everything in AI is running partly on investors money, everyone is trying to buy a monopoly and insurmountable lead and some way to lock people into a specific model and ecosystem, but so far that hasn't happened (except for people who voluntarily lock themselves into a specific ecosystem, but even in those cases, it's usually easy to get the AI to help move to another, there are no truly unique features in AI that at least one, and probably three or four, other players don't also offer).
Re: Apple Silicon costs more than OpenRouter
#273Earlier quoted context omitted.
Software that is sold as a service and requires ongoing maintenance like running in the cloud (and people to keep it running in the cloud) is opex not capex. Google Search is most definitely opex.
There was some accounting changes passed during the Biden administration’s funding bill that made it capex for a few years. It’s been rolled back in the Trump administration’s bill and I suspect no one will ever let it happen again.
https://kpmg.com/kpmg-us/content/dam/kpmg/pdf/2023/tcja-chan...
Re: Apple Silicon costs more than OpenRouter
#274Earlier quoted context omitted.
That seems unlikely. There are many providers for open models on openrouter. It seems unlikely that they are throwing money away for each token they sell. Also, there a good technical reasons for inference being much more efficient at scale.
The providers on OpenRouter serving open models aren't "throwing money away", agreed. But that's not the point I'm making. (or, it kind of is, but it's more high level than that). They're running spot and preemptible GPU instances (60-80% cheaper than on-demand), paying wholesale industrial electricity rates, and running at multi-tenant utilisation densities that make your MacBook look like a bonfire. Of course they'…
I think the question in terms of throwing money away isn't the inference layer: it's whether the companies training open models will be able to financially keep doing so. How long will Moonshot keep releasing future Kimi models? I think there's an interesting wedge they're exploring with being basically a base-model-trainer-as-a-service, i.e. selling rights to Fireworks to sell finetuning services to the Cursors of the world, but it's entirely possible it doesn't pan out.
That being said, Nvidia seems willing to step up to being the base model trainer of last resort via the Nemotron family of open models, since it helps sell more of their hardware — similar to their investments in the CUDA stack to sell hardware (unsurprisingly, Nemotron is designed to run most efficiently on Nvidia hardware, e.g. native NVFP4). So I suspect there will continue to be a pretty good market here.
Re: Apple Silicon costs more than OpenRouter
#275Earlier quoted context omitted.
Yea this; it’s the same reason why mortgaging is cheaper than renting
Except one day the hype will catch up to reality that was always true, people will realize their $20,000 Mac is has less utility as a "way to learn AI" than some kids 3090 fortnite machine, and it'll be back to below MSRP.
Re: Apple Silicon costs more than OpenRouter
#276This isn't a good analysis, and it's because it keeps rounding everything up. He rounds up the cost of electricity by 10%. He has a range of power use, takes the high end (which is 2x the low end) and multiplies it by the inflated electricity cost. But then they talk about using a newly purchased Mac to do the inference, running at full capacity, 24/7. Why would you do that? Apple silicon is fast but the author point…
Re: Apple Silicon costs more than OpenRouter
#277Re: Apple Silicon costs more than OpenRouter
#278Earlier quoted context omitted.
But once all that is done you still own a Mac in one case, and you don’t in the other, correct?
Even at just the electricity cost openrouter will be both 1) Roughly break-even to a little bit cheaper per token cost 2) Much, much, faster So the cost of the mac barely even matters, it's just an extra cost beyond. Sure, data center providers can pay lower rates. The point of this article is that LLMs at home really don't make a ton of sense, unless you are willing to pay through the nose for privacy. There is abso…
Because one day they'll send you an email informing you the new rate is $1.50, and if you missed the email, that's not their problem.
Re: Apple Silicon costs more than OpenRouter
#279Earlier quoted context omitted.
They're responding to the people doing things like buying the most expensive Mac they can find specifically to do local inference for their AI agents. Some do it to have control over their ability to use AI. Some do it because they think it will be cheaper to not have to pay a SaaS to generate tokens for them. But for those interested in the latter case, it seems like it's not actually cheaper after all, at least at…
You also have control over your costs. It is reasonable to assume that tokens will cost significantly more in the near to medium future as the market consolidates and subsidies decline.
Re: Apple Silicon costs more than OpenRouter
#280This isn't a good analysis, and it's because it keeps rounding everything up. He rounds up the cost of electricity by 10%. He has a range of power use, takes the high end (which is 2x the low end) and multiplies it by the inflated electricity cost. But then they talk about using a newly purchased Mac to do the inference, running at full capacity, 24/7. Why would you do that? Apple silicon is fast but the author point…
Not sure where 40 tokens per second is coming from. I’ve seen 95-100 tokens per second on M5 Max 128GB running Gemma 4 31B. I’ve done experiments where it is faster than Claude Opus 4.5 for the same prompts.