Live data from Hacker News

Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

neomindlabs.com

91–100 of 227 posts

Re: Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

#91
post #44

Some ppl don't like to hear it. But I would assume that token costs when using an inference provider are cheaper than electricity of using locally. If we just take into account output token generation for simplicity. With 5tps u get 18k tokens an hour. That would costs around 0.005USD from an inference provider. I estimate that the server consumes probably around 500W during inference. In Germany where 1kwh cost arou…

I don't pay anywhere near 0.30usd in the US - I pay half that off peak and can buy 1000$ worth of batteries to load up on super off peak (0.11usd). Also the inference providers are fighting over market share with huge debt loads so they are definitely going to go up in price.

Inference costs will go down massively once they use the upcoming GPUs. I estimated that a model like GLM5.2 will be around 0.03USD/M output tokens in 2 years when the Feynman GPUs will be available in 2028. And this did not even consider architectural efficiency improvements. In mid 2027 we will already see a 10x reduction once everyone has switched to the Ruby architecture.

It will be feasible for everyone to have 20 different agents running at all times. A new world is coming

Re: Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

#92
post #68

Earlier quoted context omitted.

It's still the same thing, you can ask it to do a full on report give explanation and details be thorough and then go do something else, another task a lunch break whatever and it will be done when you're back

How do you maintain a flow state during a lunch break? I'm looping with Claude on a scale of minutes. While you're waiting, I'm iterating.

[deleted]

Re: Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

#93
post #87

Earlier quoted context omitted.

> Once you've finished reading the next article on HN about a 5 tps Xeon, your task will be complete. If I spend 10 minutes reading an article, that would only generate 3000 tokens. That’s not counting the prompt processing time. We have very different expectations for LLMs if your tasks only take a couple thousand tokens and you’re happy waiting 10 minutes for it. > Yes, with top-tier GPU farms you can hit hundreds…

You could suspend it to ram, and only wake it up on request, it takes 2 seconds on my box.

It’s not a cost savings relative to paying API prices even if you’re suspending it.

This is an option if you must run local inference, you’re not sensitive to speed, and the budget is low.

It’s not going to be cheaper than paying API prices for the model though.

Re: Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

#94
post #21

I have a prediction. By the mid of 2027, we will have >200B MoE models running on basic consumer hardware. I am running Qwen3.6-35B-A3B locally on my 16GB mac with 7-9 tokens/second. Link - https://github.com/deepanwadhwa/samosa-chat This is a GPT4 level model running locally with a decent speed on a 16gb ram macbook air.

Downloading now just 'cause the repo name

Re: Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

#95
post #21

I have a prediction. By the mid of 2027, we will have >200B MoE models running on basic consumer hardware. I am running Qwen3.6-35B-A3B locally on my 16GB mac with 7-9 tokens/second. Link - https://github.com/deepanwadhwa/samosa-chat This is a GPT4 level model running locally with a decent speed on a 16gb ram macbook air.

I tried Qwen3.6-35B-A3B, but it couldn't generate a 50-100 line Clojure file without having broken parens mismatches. I know Clojure isn't super popular, but the syntax is pretty simple and the frontier models do fine with it.

try q8, check your parameters. qwen3.6-35b-a3b should definitely be able to do so with no issues at all.

Re: Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

#96
post #44

Some ppl don't like to hear it. But I would assume that token costs when using an inference provider are cheaper than electricity of using locally. If we just take into account output token generation for simplicity. With 5tps u get 18k tokens an hour. That would costs around 0.005USD from an inference provider. I estimate that the server consumes probably around 500W during inference. In Germany where 1kwh cost arou…

don't care, and yeah i don't like to hear it. we don't run local because it's cheaper money wise. we do it for freedom, for privacy and having option makes it cheaper in the long run. if there was no local options, your cloud model would cost much more!

Re: Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

#97
post #15

Earlier quoted context omitted.

This reads as pretty clearly AI-generated text, which is against HN guidelines.

The PR? He said it was AI in the comment you replied to... I don't think the post itself reads like AI at all, but that's just me.

The post is absolutely LLM-generated. “Punchy” short sentences, “… has quietly come to mean …”, “The optimized paths weren’t there to execute.”

Re: Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

#98

A dual Xeon of this era is probably pulling 300W or more when loaded. At national average electricity prices, that’s $1.35 per day. More during the summer if you have to cool the space. If you run it 24/7 and ignore prompt processing time (not a good assumption at all) it would get around 400,000 tokens in a day. That’s about $0.30 per million output tokens. Coincidentally, that’s the same price for this model on Ope…

It gets better in the cooler months when heating is running in a home :)

Re: Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

#99
post #68

Earlier quoted context omitted.

It's still the same thing, you can ask it to do a full on report give explanation and details be thorough and then go do something else, another task a lunch break whatever and it will be done when you're back

How do you maintain a flow state during a lunch break? I'm looping with Claude on a scale of minutes. While you're waiting, I'm iterating.

This is like comparing a hammer to a screwdriver and feeling smug because you can hammer nails faster than someone else can drive screws.

These are fundamentally different tools for entirely different applications. They only look similar to people who don't understand the tools or their purpose.

Re: Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

#100
post #83

Earlier quoted context omitted.

I think, may be actually wrong, that most of us do not consider running a model locally a way to save money. It is a way not to spread personal info around.

Anyone running LLMs at home will come to that realization quickly, if they’re looking at their power bills. Even feeling the heat output of a computer running at 100% in your office makes it clear. I was responding to a lot of the comments saying this was a reasonable way to avoid paying for tokens or subscriptions. I don’t want anyone getting the wrong idea that this is a way to save money if that’s their priority.

> Even feeling the heat output of a computer running at 100% in your office makes it clear.

What does it make clear? That I can replace the space heater my wife runs 9 out of 12 months of the year with a home server? And effectively get $0.00 per token during those times?

In houses running A/C year round, sure there'd be some impact, but in all the places running heat, doesn't seem that it'd move the needle on power bills.

There are startups whose entire business model is "cloud server as a home space heater" (aka "data furnace") ...

Post reply on HN