Live data from Hacker News

Apple Silicon costs more than OpenRouter

williamangel.net

261–270 of 322 posts

Re: Apple Silicon costs more than OpenRouter

#262

Earlier quoted context omitted.

Rounding everything down in the most optimistic setting got me to $0.40 per million tokens, and openrouter has the same model at $.38/mtok.

What is it with AI SaaS naming themselves "openxyz" when there is 0% open about them?

It's how marketing works. If something is a problem they have to loudly claim to have fixed it. Look around the economy and you'll see lots of it. "Healthy" (high sugar) muesli bars, clean-diesel, surveillance wrapped up as keeping us safe. The modus operandi of marketing is to change minds about self evident things otherwise what is the point?

Re: Apple Silicon costs more than OpenRouter

#264
post #179

I think that the main flaw in the reasoning is assuming that cost of token will stay the same over the years. Chances are that token prices will go down, but chances also are that the AI bubble pops and all of a sudden all these companies will either have to make a buck out of the inference or go bankrupt. Getting your own hardware just grants you stable pricing.

if all those companies go bankrupt, you could also buy their hardware on the cheap.

its almost guaranteed imo that as the model quality evens out, inference cost will drop towards 0 at insane speeds, given how well llama works as an asic.

Re: Apple Silicon costs more than OpenRouter

#265
post #59

Frontier AI companies are selling at a loss. Excusing everything else that u/bastawhiz said[0]; the obvious fact here is that Claude, OpenAI, Gemini et al. are quite literally burning through 100's of billions of dollars and selling it back to you for pennies on the dollar in the hopes that they get to be the only one left. If I spend $10 growing Oranges and sell them to you for $1; then of course it's more expensive…

Do you have a proof? Anthropic’s CEO said they Are profitable. Same with OpenAI.

do you have proof? Taking these guys at face value is not wise

Re: Apple Silicon costs more than OpenRouter

#266
This is not surprising at all. The biggest benefit of cloud model in terms of energy efficiency is that when running more than 1 requests, the power consumption of said GPU roughly stayed the same. The more concurrency requests the server can handle, the less power each request consume. The server GPU is already likely more energy efficient than local GPU, concurrency make the cost structure unbeatable by local hardware. It is generally assumed the local hardware only run 1 request, but if the local engine is meant to serve a small business with meaningful concurrency, the economy might still work out.

Re: Apple Silicon costs more than OpenRouter

#268
For me, the value in local inference is getting your hands dirty and goofing around. That is to say, learning.

So we shouldn’t be comparing it to the cost of open router api access at all, we should be comparing it to the cost of a 4 credit university course.

Re: Apple Silicon costs more than OpenRouter

#269

This isn't a good analysis, and it's because it keeps rounding everything up. He rounds up the cost of electricity by 10%. He has a range of power use, takes the high end (which is 2x the low end) and multiplies it by the inflated electricity cost. But then they talk about using a newly purchased Mac to do the inference, running at full capacity, 24/7. Why would you do that? Apple silicon is fast but the author point…

Boss, I make 16.50 per hour, say 15, I work 36 hours, say 35, say 500 per week, say 4 weeks per month, that's only about 2000! Don't you agree I need a raise?

Re: Apple Silicon costs more than OpenRouter

#270

If you want a good dense model, use qwen3.6 27B instead, speed will be up, and if you don't take my word for it being smarter, take openrouter's prices of it against the bigger, slower and less memory-efficient gemma do the talking. If you want a faster model, go for qwen3.6 35B (or gemma 4 26B if gemma models perform better for your tasks). There is a reason why people (myself included) haven't shut up about those t…

I'm interested in how you evaluate quantized models against each other; haven't found a benchmark I love for that. I love this example about 27B debugging. I've seen similar success after I got a Mac with 4x memory; and Qwen 35B A3B all of a sudden is doing a great job (the 9B on my laptop wasn't great to say the least).
Post reply on HN