Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
31–40 of 171 posts
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#32Everything runs on a π if you quantize it enough! I'm curious about the applications though. Do people randomly buy 4xRPi5s that they can now dedicate to running LLMs?
I'd love to hook my development tools into a fully-local LLM. The question is context window and cost. If the context window isn't big enough, it won't be helpful for me. I'm not gonna drop $500 on RPis unless I know it'll be worth the money. I could try getting my employer to pay for it, but I'll probably have a much easier time convincing them to pay for Claude or whatever.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#33Everything runs on a π if you quantize it enough! I'm curious about the applications though. Do people randomly buy 4xRPi5s that they can now dedicate to running LLMs?
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#34Everything runs on a π if you quantize it enough! I'm curious about the applications though. Do people randomly buy 4xRPi5s that they can now dedicate to running LLMs?
I'd love to hook my development tools into a fully-local LLM. The question is context window and cost. If the context window isn't big enough, it won't be helpful for me. I'm not gonna drop $500 on RPis unless I know it'll be worth the money. I could try getting my employer to pay for it, but I'll probably have a much easier time convincing them to pay for Claude or whatever.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#35Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#36Earlier quoted context omitted.
13/s is not slow. Q4 is not bad. The models that run on phones are never 30B or anywhere close to that.
It is very slow and totally unimpressive. 5060Ti ($430 new) would do over 60, even more in batched mode. 4x RPi 5 are $550 new.
Yes, I'm joking.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#37Everything runs on a π if you quantize it enough! I'm curious about the applications though. Do people randomly buy 4xRPi5s that they can now dedicate to running LLMs?
I'd love to hook my development tools into a fully-local LLM. The question is context window and cost. If the context window isn't big enough, it won't be helpful for me. I'm not gonna drop $500 on RPis unless I know it'll be worth the money. I could try getting my employer to pay for it, but I'll probably have a much easier time convincing them to pay for Claude or whatever.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#38Everything runs on a π if you quantize it enough! I'm curious about the applications though. Do people randomly buy 4xRPi5s that they can now dedicate to running LLMs?
I have clusters of over a thousand raspberry pi’s that have generally 75% of their compute and 80% of their memory that is completely unused.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#39Earlier quoted context omitted.
[flagged]
this is very pessimistic take there are lot of bad people on internet too, does that make internet is a mistake ??? Noo, the people are not the tool
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#40This is really impressive. If we can get this down to a single Raspberry Pi, then we have crazy embedded toys and tools. Locally, at the edge, with no internet connection. Kids will be growing up with toys that talk to them and remember their stories. We're living in the sci-fi future. This was unthinkable ten years ago.