Live data from Hacker News

Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5

github.com

81–90 of 171 posts

Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5

#81
post #50

Earlier quoted context omitted.

> I'd love to hook my development tools into a fully-local LLM. Karpathy said in his recent talk, on the topic of AI developer-assistants: don't bother with less capable models. So ... using an rpi is probably not what you want.

It's a tough thing, I'm a solo dev supporting ~all at high quality. I cannot imagine using anything other than $X[1] at the leading edge. Why not have the very best? Karpathy elides he is an individual. We expect to find a distribution of individuals, such that a nontrivial # of them are fine with 5-10% off the leading edge performance. Why? At least for free as in beer. At most, concerns about connectivity, IP right…

Today's qwen3 30b is about as good as last year's state of the art. For me that's more than good enough. Many tasks don't require the best of the best either.

Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5

#82
post #50

Earlier quoted context omitted.

I'd love to hook my development tools into a fully-local LLM. The question is context window and cost. If the context window isn't big enough, it won't be helpful for me. I'm not gonna drop $500 on RPis unless I know it'll be worth the money. I could try getting my employer to pay for it, but I'll probably have a much easier time convincing them to pay for Claude or whatever.

> I'd love to hook my development tools into a fully-local LLM. Karpathy said in his recent talk, on the topic of AI developer-assistants: don't bother with less capable models. So ... using an rpi is probably not what you want.

Mind linking to "his recent talk"? There's a lot of videos of him so it's a bit difficult to find what's most recent.

Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5

#83

Earlier quoted context omitted.

That’s an interesting setup. What are you doing with that sort of cluster?

99.9% of enthusiast/hobbyist clusters like this are exclusively used for blinkenlights

Blinkenlights are an admirable pursuit

Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5

#84
post #21

Earlier quoted context omitted.

I think it is all well and good, but the most affordable option is probably still to buy a used MacBook with 16/32 or 64 GB (depending on the budget) unified memory and install Asahi Linux for tinkering. Graphics cards with decent amount of memory are still massively overpriced (even used), big, noisy and draw a lot of energy.

It just came to my attention that the 2021 M1 Max 64gb is less than $1500 used. That’s 64gb of unified memory at regular laptop prices, so I think people will be well equipped with AI laptops rather soon. Apple really is #2 and probably could be #1 in AI consumer hardware.

M1 doesn't exactly have stellar memory bandwidth for this day and age though

Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5

#85
post #11

Earlier quoted context omitted.

It's not unrealistically pessimistic. We're already seeing research showing the negative effects, as well as seeing routine psychosis stories. Think about the ways that LLMs interact. The constant barrage of positive responses "brilliant observation" etc. That's not a healthy input to your mental feedback loop. We all need responses that are grounded in reality, just like you'd get from other human beings. Think abou…

The irony of this is that Gen-Z have been mollycoddled with praise by their parents and modern life, we give medals for participation, or runners up prizes for losing. We tell people when they've failed at something they did their best and that's what matters. We validate their upset feelings if they're insulted by free speech that goes against their beliefs. This is exactly what is happening with sycophantic LLMs, t…

You sound very old man yelling at cloud. And the winner takes all is so American.

And no discrimination against lgbt etc under the guise of free speech is not ok.

Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5

#87
post #63
post #62

Earlier quoted context omitted.

How do you evaluate this except for anecdote and how do we know your experience isn't due to how you use them?

You can evaluate it as anecdote. How do I know you have the level of experience necessary to spot these kinds of problems as they arise? How do I know you're not just another AI booster with financial stake poisoning the discussion? We could go back and forth on this all day.

you got very defensive. it was a useful question - they were asking in terms of using a local LLM, so at best they might be in the business of selling raspberry pis, not proprietary LLMs.

Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5

#89
post #71

Earlier quoted context omitted.

GPT OSS 20B is smart enough but the context window is tiny with enough files. Wonder if you can make a dumber model with a massive context window thats a middleman to GPT.

Matches my experience.

Just have it open a new context window, the other thing I wanted to try is to make a LoRa but im not sure how that works properly, it suggested a whole other model but it wasnt a pleasant experience since it’s not as obvious as diffusion models for images.
Post reply on HN