Earlier quoted context omitted.
> I'd love to hook my development tools into a fully-local LLM. Karpathy said in his recent talk, on the topic of AI developer-assistants: don't bother with less capable models. So ... using an rpi is probably not what you want.
It's a tough thing, I'm a solo dev supporting ~all at high quality. I cannot imagine using anything other than $X[1] at the leading edge. Why not have the very best? Karpathy elides he is an individual. We expect to find a distribution of individuals, such that a nontrivial # of them are fine with 5-10% off the leading edge performance. Why? At least for free as in beer. At most, concerns about connectivity, IP right…
Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
81–90 of 171 posts
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#82Earlier quoted context omitted.
I'd love to hook my development tools into a fully-local LLM. The question is context window and cost. If the context window isn't big enough, it won't be helpful for me. I'm not gonna drop $500 on RPis unless I know it'll be worth the money. I could try getting my employer to pay for it, but I'll probably have a much easier time convincing them to pay for Claude or whatever.
> I'd love to hook my development tools into a fully-local LLM. Karpathy said in his recent talk, on the topic of AI developer-assistants: don't bother with less capable models. So ... using an rpi is probably not what you want.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#83Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#84Earlier quoted context omitted.
I think it is all well and good, but the most affordable option is probably still to buy a used MacBook with 16/32 or 64 GB (depending on the budget) unified memory and install Asahi Linux for tinkering. Graphics cards with decent amount of memory are still massively overpriced (even used), big, noisy and draw a lot of energy.
It just came to my attention that the 2021 M1 Max 64gb is less than $1500 used. That’s 64gb of unified memory at regular laptop prices, so I think people will be well equipped with AI laptops rather soon. Apple really is #2 and probably could be #1 in AI consumer hardware.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#85Earlier quoted context omitted.
It's not unrealistically pessimistic. We're already seeing research showing the negative effects, as well as seeing routine psychosis stories. Think about the ways that LLMs interact. The constant barrage of positive responses "brilliant observation" etc. That's not a healthy input to your mental feedback loop. We all need responses that are grounded in reality, just like you'd get from other human beings. Think abou…
The irony of this is that Gen-Z have been mollycoddled with praise by their parents and modern life, we give medals for participation, or runners up prizes for losing. We tell people when they've failed at something they did their best and that's what matters. We validate their upset feelings if they're insulted by free speech that goes against their beliefs. This is exactly what is happening with sycophantic LLMs, t…
And no discrimination against lgbt etc under the guise of free speech is not ok.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#86Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#87Earlier quoted context omitted.
How do you evaluate this except for anecdote and how do we know your experience isn't due to how you use them?
You can evaluate it as anecdote. How do I know you have the level of experience necessary to spot these kinds of problems as they arise? How do I know you're not just another AI booster with financial stake poisoning the discussion? We could go back and forth on this all day.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#88So would 40x RPi 5 get 130 token/s?
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#89Earlier quoted context omitted.
GPT OSS 20B is smart enough but the context window is tiny with enough files. Wonder if you can make a dumber model with a massive context window thats a middleman to GPT.
Matches my experience.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#90Everything runs on a π if you quantize it enough! I'm curious about the applications though. Do people randomly buy 4xRPi5s that they can now dedicate to running LLMs?
though at what quality?