Deepseek R1 Distill 8B Q40 on 4 x Raspberry Pi 5
1–10 of 162 posts
Re: Deepseek R1 Distill 8B Q40 on 4 x Raspberry Pi 5
#2This is the modern Beowulf cluster.
Re: Deepseek R1 Distill 8B Q40 on 4 x Raspberry Pi 5
#3As always, take those t/s stats with a huge boulder of salt. The demo shows a question "solved" in < 500 tokens. Still amazing that it's possible, but you'll get nowhere near those speeds when dealing with real-world problems at real-world useful context lengths for "thinking" models (8-16k tokens). Even epyc's with lots of channels go down to 2-4 t/s after ~4096 context length.
Re: Deepseek R1 Distill 8B Q40 on 4 x Raspberry Pi 5
#4This continues the pattern of all other announcements of running 'Deepseek R1' on raspberry pi - that they are running llama (or qwen), modified by deepseek's distillation technique.
Re: Deepseek R1 Distill 8B Q40 on 4 x Raspberry Pi 5
#5Okay but does any one actually _want_ a reasoning model at such low tok/sec speeds?!
Re: Deepseek R1 Distill 8B Q40 on 4 x Raspberry Pi 5
#6That’s the real future
Re: Deepseek R1 Distill 8B Q40 on 4 x Raspberry Pi 5
#7This is the modern Beowulf cluster.
But can it run Crysis?
Re: Deepseek R1 Distill 8B Q40 on 4 x Raspberry Pi 5
#8Okay but does any one actually _want_ a reasoning model at such low tok/sec speeds?!
Lots of use cases don’t require low latency. Background work for agents. CI jobs. Other stuff I haven’t thought of.
Re: Deepseek R1 Distill 8B Q40 on 4 x Raspberry Pi 5
#9Okay but does any one actually _want_ a reasoning model at such low tok/sec speeds?!
No, but the alternative in some places is no reasoning model. Just like people don't want old cell phones / new phones with old chips - but often that's all that is affordable in some places.
If we can get something working, then improving it will come.
Re: Deepseek R1 Distill 8B Q40 on 4 x Raspberry Pi 5
#10This continues the pattern of all other announcements of running 'Deepseek R1' on raspberry pi - that they are running llama (or qwen), modified by deepseek's distillation technique.
Yeah. People looking for “Smaller DeepSeek” are looking for the quantized models, which are still quite large.