Live data from Hacker News

Deepseek R1 Distill 8B Q40 on 4 x Raspberry Pi 5

github.com

1–10 of 162 posts

Re: Deepseek R1 Distill 8B Q40 on 4 x Raspberry Pi 5

#3
As always, take those t/s stats with a huge boulder of salt. The demo shows a question "solved" in < 500 tokens. Still amazing that it's possible, but you'll get nowhere near those speeds when dealing with real-world problems at real-world useful context lengths for "thinking" models (8-16k tokens). Even epyc's with lots of channels go down to 2-4 t/s after ~4096 context length.

Re: Deepseek R1 Distill 8B Q40 on 4 x Raspberry Pi 5

#9
post #5

Okay but does any one actually _want_ a reasoning model at such low tok/sec speeds?!

No, but the alternative in some places is no reasoning model. Just like people don't want old cell phones / new phones with old chips - but often that's all that is affordable in some places.

If we can get something working, then improving it will come.

Re: Deepseek R1 Distill 8B Q40 on 4 x Raspberry Pi 5

#10
post #4

This continues the pattern of all other announcements of running 'Deepseek R1' on raspberry pi - that they are running llama (or qwen), modified by deepseek's distillation technique.

Yeah. People looking for “Smaller DeepSeek” are looking for the quantized models, which are still quite large.

https://unsloth.ai/blog/deepseekr1-dynamic

Post reply on HN