Live data from Hacker News

Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

neomindlabs.com

101–110 of 227 posts

Re: Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

#101
post #27

Earlier quoted context omitted.

> on a decent speed But you said 7-9 tokens/second , that's not a decent speed. I'm not an expert by all means but in my local experiments, less than 12 to 16 tps is too slow.

I think that any workflow that requires the user to stare at the tokens being generated live is using it wrong. Delegate, don't stare! https://mikeveerman.github.io/tokenspeed/?rate=10&mode=text You think of an idea that you want to have the LLM process, queue it up, and go back to what you were doing. Once you've finished reading the next article on HN about a 5 tps Xeon, your task will be complete. It's kind of lik…

Except often queued agentic flows must be checked in on. Or to use the comparison, 3D printers are not immune to making spaghetti all night when something goes wrong. (I’m not a 3d printing expert so maybe that is solved now)

It is common for agents to just stop because overload or some API error hijinks.

Or you get a TUI question that is blocking.

In general you’re right though, staring at tokens from agentic is not time well spent.

Some of these I’ve built custom harness around in iterm2 though.

Re: Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

#103
post #59

Earlier quoted context omitted.

You are comparing a 35B models to a 635B+ frontier model, of course thats not even close

I'm not lamenting that they aren't close, I'm saying Qwen will frequently output code that isn't even syntactically correct, even when the syntax is simple. Which makes it unusable for coding.

It really depends on the language, popular languages work pretty good

Re: Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

#104
post #89
post #21

I have a prediction. By the mid of 2027, we will have >200B MoE models running on basic consumer hardware. I am running Qwen3.6-35B-A3B locally on my 16GB mac with 7-9 tokens/second. Link - https://github.com/deepanwadhwa/samosa-chat This is a GPT4 level model running locally with a decent speed on a 16gb ram macbook air.

By early 2028, major players like Intel, AMD, QC will ship accelerators in consumer laptops capable of running ~1T MoE models at ~100 tok/s

Unless there are major improvements to how much hardware it takes to run a 1T model, this is deeply unrealistic. First because why release hardware that puts your biggest customers (data centers) out of business. Second because as I understand it the data centers have bought up all the high end chip production capacity for at least the next year and unless the bubble pops that'll continue for a while.

Re: Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

#105
post #27

Earlier quoted context omitted.

> on a decent speed But you said 7-9 tokens/second , that's not a decent speed. I'm not an expert by all means but in my local experiments, less than 12 to 16 tps is too slow.

I think that any workflow that requires the user to stare at the tokens being generated live is using it wrong. Delegate, don't stare! https://mikeveerman.github.io/tokenspeed/?rate=10&mode=text You think of an idea that you want to have the LLM process, queue it up, and go back to what you were doing. Once you've finished reading the next article on HN about a 5 tps Xeon, your task will be complete. It's kind of lik…

Filament snaps at 1am and then you have to run print again. 10 hours turn into many days potentially.

I watch tokens to see if it goes in right direction. If model goes off the rails, then it is time to stop and adjust prompt.

Re: Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

#106
post #44

Some ppl don't like to hear it. But I would assume that token costs when using an inference provider are cheaper than electricity of using locally. If we just take into account output token generation for simplicity. With 5tps u get 18k tokens an hour. That would costs around 0.005USD from an inference provider. I estimate that the server consumes probably around 500W during inference. In Germany where 1kwh cost arou…

> Some ppl don't like to hear it. But I would assume that token costs when using an inference provider are cheaper than electricity of using locally.

Maybe, but for how long? Prices keep going up, and every new model eats more and more tokens...

Re: Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

#107
post #89
post #21

I have a prediction. By the mid of 2027, we will have >200B MoE models running on basic consumer hardware. I am running Qwen3.6-35B-A3B locally on my 16GB mac with 7-9 tokens/second. Link - https://github.com/deepanwadhwa/samosa-chat This is a GPT4 level model running locally with a decent speed on a 16gb ram macbook air.

By early 2028, major players like Intel, AMD, QC will ship accelerators in consumer laptops capable of running ~1T MoE models at ~100 tok/s

Literally the only way this is going to happen is if aliens come to earth and gift us some amazing technology.

Re: Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

#109
post #21

I have a prediction. By the mid of 2027, we will have >200B MoE models running on basic consumer hardware. I am running Qwen3.6-35B-A3B locally on my 16GB mac with 7-9 tokens/second. Link - https://github.com/deepanwadhwa/samosa-chat This is a GPT4 level model running locally with a decent speed on a 16gb ram macbook air.

> I have a prediction. By the mid of 2027, we will have >200B MoE models running on basic consumer hardware. This prediction alone isn’t useful at all without a bound on speed and maybe quantization. You can already run >200B MoE models on basic consumer hardware by picking a low bpw quantization and then streaming the experts from SSD. There have been a lot of proof of concept demos, but nobody uses them because the…

>>Yes, it’s technically running, but not in a way that would be useful by normal LLM standards.

What are the LLM standards?

Do you know how many people use perplexity? I know many people who are not software engineers or tech workers and have a LLM subscription for rewriting their stuff (non-native english speakers) in english. There are many use cases for running good models locally. Maybe not for you, but someone might find this beneficial.

Re: Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

#110
post #89

Earlier quoted context omitted.

By early 2028, major players like Intel, AMD, QC will ship accelerators in consumer laptops capable of running ~1T MoE models at ~100 tok/s

Unless there are major improvements to how much hardware it takes to run a 1T model, this is deeply unrealistic. First because why release hardware that puts your biggest customers (data centers) out of business. Second because as I understand it the data centers have bought up all the high end chip production capacity for at least the next year and unless the bubble pops that'll continue for a while.

Because for the company that will actually do it, their biggest customers aren’t data centers they are iPhone owners.
Post reply on HN