Earlier quoted context omitted.
Model intelligence should be part of your equation as well, unless you love loads and loads of hidden technical debt and context-eating, unnecessarily complex abstractions
GPT OSS 20B is smart enough but the context window is tiny with enough files. Wonder if you can make a dumber model with a massive context window thats a middleman to GPT.
Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
71–80 of 171 posts
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#72Everything runs on a π if you quantize it enough! I'm curious about the applications though. Do people randomly buy 4xRPi5s that they can now dedicate to running LLMs?
I'd love to hook my development tools into a fully-local LLM. The question is context window and cost. If the context window isn't big enough, it won't be helpful for me. I'm not gonna drop $500 on RPis unless I know it'll be worth the money. I could try getting my employer to pay for it, but I'll probably have a much easier time convincing them to pay for Claude or whatever.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#73Very impressive numbers.. wonder how this would scale on 4 relatively modern desktop PCs, like say something akin to a i5 8th Gen Lenovo ThinkCentre, these can be had for very cheap. But like @geerlingguy indicates - we need model compatibility to go up up up! As an example it would amazing to see something like fastsdcpu run distributed to democratize accessibility-to/practicality-of image gen models for people with…
I think it is all well and good, but the most affordable option is probably still to buy a used MacBook with 16/32 or 64 GB (depending on the budget) unified memory and install Asahi Linux for tinkering. Graphics cards with decent amount of memory are still massively overpriced (even used), big, noisy and draw a lot of energy.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#74Earlier quoted context omitted.
People don’t know what they want yet, you have to show it to them. Getting the hardware out is part of it, but you are right, we’re missing the killer apps at the moment. The very need for privacy with AI will make personal hardware important no matter what.
Two main factors are holding back the "killer app" for AI. Fix hallucinations and make agents more deterministic. Once these are in place, people will love AI when it can make them money somehow.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#75Earlier quoted context omitted.
> there are lot of bad people on internet too, does that make internet is a mistake ??? Yes. People write and say “the Internet was a mistake” all the time, and some are joking, but a lot of us aren’t.
are you going to give up knife too because some people use it for crime????
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#76Earlier quoted context omitted.
People don’t know what they want yet, you have to show it to them. Getting the hardware out is part of it, but you are right, we’re missing the killer apps at the moment. The very need for privacy with AI will make personal hardware important no matter what.
Two main factors are holding back the "killer app" for AI. Fix hallucinations and make agents more deterministic. Once these are in place, people will love AI when it can make them money somehow.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#77Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#78Earlier quoted context omitted.
I think it is all well and good, but the most affordable option is probably still to buy a used MacBook with 16/32 or 64 GB (depending on the budget) unified memory and install Asahi Linux for tinkering. Graphics cards with decent amount of memory are still massively overpriced (even used), big, noisy and draw a lot of energy.
What about AMD Ryzen AI Max+ 395 Mini PCs with upto 128GB unified memory?
Seems like at the consumer hardware level you just have to pick your poison or what one factor you care about most. Macs with a Max or Ultra chip can have good memory bandwidth but low compute, but also ultra low power consumption. Discrete GPUs have great compute and bandwidth but low to middling VRAM, and high costs and power consumption. The unified memory PCs like the Ryzen AI Max and the Nvidia DGX deliver middling compute, higher VRAMs, and terrible memory bandwidth.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#79This is really impressive. If we can get this down to a single Raspberry Pi, then we have crazy embedded toys and tools. Locally, at the edge, with no internet connection. Kids will be growing up with toys that talk to them and remember their stories. We're living in the sci-fi future. This was unthinkable ten years ago.
[flagged]
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#80At this speed this is only suitable for time insensitive applications..