Earlier quoted context omitted.
It's a tough thing, I'm a solo dev supporting ~all at high quality. I cannot imagine using anything other than $X[1] at the leading edge. Why not have the very best? Karpathy elides he is an individual. We expect to find a distribution of individuals, such that a nontrivial # of them are fine with 5-10% off the leading edge performance. Why? At least for free as in beer. At most, concerns about connectivity, IP right…
Today's qwen3 30b is about as good as last year's state of the art. For me that's more than good enough. Many tasks don't require the best of the best either.
Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
151–160 of 171 posts
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#152Earlier quoted context omitted.
I'd love to hear more about what you're running, and on what hardware. Also, what is your use case? Thanks!
So I am running Ollama on Windows using an 10700k and 3080ti. I'm using models like Qwen3-coder (4/8b) and 2.5-coder 15b, Llama 3 instruct, etc. These models are very fast on my machine (~25-100 tokens per second depending on model) My use case is custom software that I build and host that leverages LLMs for example for domotica where I use my Apple watch shortcuts to issue commands. I also created a VS2022 extension…
Have a great week.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#153Earlier quoted context omitted.
> Karpathy said in his recent talk, on the topic of AI developer-assistants: don't bother with less capable models. Interesting because he also said the future is small "cognitive core" models: > a few billion param model that maximally sacrifices encyclopedic knowledge for capability. It lives always-on and by default on every computer as the kernel of LLM personal computing. https://xcancel.com/karpathy/status/1938…
It's not at all trivial to build a "small but highly capable" model. Sacrificing world knowledge is something that can be done, but only to an extent, and that isn't a silver bullet. For an LLM, size is a virtue - the larger a model is, the more intelligent it is, all other things equal - and even aggressive distillation only gets you this far. Maybe with significantly better post-training, a lot of distillation from…
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#154Earlier quoted context omitted.
M1 Max with 64GB has 400GB/s memory bandwidth. You have to get into the highest 16-core M4 Max configurations to begin pulling away from that number.
Oh sorry I thought it was only about 100. I'd read that before but I must have remembered incorrectly. 400 is indeed very serviceable.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#155Earlier quoted context omitted.
Not just LPDDR5, but LPDDR5X-8000 on a 256-bit bus. The 40 CU of RDNA 3.5 is nice, but it's less raw compute than e.g. a desktop 4060 Ti dGPU. The memory is fast, 200+ GB/s real-world read and write (the AIDA64 thread about limited read speeds is misleading, this is what the CPU is able to see, the way the memory controller is configured, but GPU tooling reveals full 200+ GB/s read and write). Though you can only all…
The 128 GB Strix Halo system was tempting me, but I think I'm going to hold out for the Medusa Point memory bandwidth gains to expand my cluster setup. I have a Mac Mini M4 Pro 64GB that does quite well with inference on the Qwen3 models, but is hell on networking with my home K3s cluster, which going deeper on is half the fun of this stuff for me.
NVDIA is so greedy that doling out $500 dollars will only you get you 16gb of vram at half the speed of a M1 Max. You can get a lot more speed with more expensive NVDIA GPUs, but you won’t get anything close to a decent amount of vram for less than 700-1500 dollars (well, truly, you will not get close to 32gb even).
Makes me wonder just how much secret effort is being put in by MAG7 to strip NVDIDA of this pricing power because they are absolutely price gouging.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#156Earlier quoted context omitted.
People don’t know what they want yet, you have to show it to them. Getting the hardware out is part of it, but you are right, we’re missing the killer apps at the moment. The very need for privacy with AI will make personal hardware important no matter what.
We've shown people so many times and so forcefully that they're now actively complaining about it. It's a meme. The problem isn't getting your Killer A I App in front of eyeballs. The problem is showing something useful or necessary or wanted . AI has not yet offered the common person anything they want or need! The people have seen what you want to show them, they've been forced to try it, over and over. There is no…
Nobody wants the one-millionth meeting transcription app and the one-millionth coding agent constantly, sure.
It a developer creativity issue. I personally believe the creativity is so egregious, that if anyone were to release a killer app, the entirety of the lackluster dev community will copy it into eternity to the point where you’ll think that that’s all AI can do.
This is not a great way to start off the morning, but gosh darn it, I really hate that this profession attracted so many people that just want to make a buck.
——-
You know what was the killer app for the Wii?
Wii Sports. It sold a lot of Wiis.
You have to be creative with this AI stuff, it’s a requirement.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#157Very impressive numbers.. wonder how this would scale on 4 relatively modern desktop PCs, like say something akin to a i5 8th Gen Lenovo ThinkCentre, these can be had for very cheap. But like @geerlingguy indicates - we need model compatibility to go up up up! As an example it would amazing to see something like fastsdcpu run distributed to democratize accessibility-to/practicality-of image gen models for people with…
I think it is all well and good, but the most affordable option is probably still to buy a used MacBook with 16/32 or 64 GB (depending on the budget) unified memory and install Asahi Linux for tinkering. Graphics cards with decent amount of memory are still massively overpriced (even used), big, noisy and draw a lot of energy.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#158Earlier quoted context omitted.
You sound very old man yelling at cloud. And the winner takes all is so American. And no discrimination against lgbt etc under the guise of free speech is not ok.
Well you're wrong on all accounts of the veiled insults. Also, I've not stated LGBT, this has nothing to do with it, it's weird you'd even mention it.
I personally feel we should be way more in touch with our emotions especially when it comes to men.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#159Earlier quoted context omitted.
Not just LPDDR5, but LPDDR5X-8000 on a 256-bit bus. The 40 CU of RDNA 3.5 is nice, but it's less raw compute than e.g. a desktop 4060 Ti dGPU. The memory is fast, 200+ GB/s real-world read and write (the AIDA64 thread about limited read speeds is misleading, this is what the CPU is able to see, the way the memory controller is configured, but GPU tooling reveals full 200+ GB/s read and write). Though you can only all…
The 128 GB Strix Halo system was tempting me, but I think I'm going to hold out for the Medusa Point memory bandwidth gains to expand my cluster setup. I have a Mac Mini M4 Pro 64GB that does quite well with inference on the Qwen3 models, but is hell on networking with my home K3s cluster, which going deeper on is half the fun of this stuff for me.
I was initially thinking this way too, but I realized a 128GB Strix Halo system would make an excellent addition to my homelab / LAN even once it's no longer the star of the stable for LLM inference - i.e. I will probably get a Medusa Halo system as well once they're available. My other devices are Zen 2 (3600x) / Zen 3 (5950x) / Zen 4 (8840u), an Alder Lake N100 NUC, a Twin Lake N150 NUC, along with a few Pi's and Rockchip SBC's, so a Zen 5 system makes a nice addition to the high end of my lineup anyway. Not to mention, everything else I have maxed out at 2.5GbE. I've been looking for an excuse to upgrade my switch from 2.5GbE to 5 or 10 GbE, and the Strix Halo system I ordered was the BeeLink GTR9 Pro with dual 10GbE. Regardless of whether it's doing LLM, other gen AI inference, some extremely light ML training / light fine tuning, media transcoding, or just being yet another UPS-protected server on my LAN, there's just so much capability offered for this price and TDP point compared to everything else I have.
Apple Silicon would've been a serious competitor for me on the price/performance front, but I'm right up there with RMS in terms of ideological hostility towards proprietary kernels. I'm not totally perfect (privacy and security are a journey, not a destination), but I am at the point where I refuse to use anything running an NT or Darwin kernel.
Re: Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
#160Earlier quoted context omitted.
Do you think I am somehow bound to answer yes to this question? If so, why do you think that?
You would not admit that because you have an ego that would expose your flawed logic
I probably consider the Internet far less valuable than you do—it’d never occur to me to compare it to knives, which are enormously useful.