Live data from Hacker News

Show HN: Slotstream, run Qwen3.8-Flash-Next 4-bit on a low-memory Mac

github.com

1–2 of 2 posts

Show HN: Slotstream, run Qwen3.8-Flash-Next 4-bit on a low-memory Mac

#1
I built slotstream, a way to run Qwen3.8-Flash-Next 4-bit on a low-memory mac starting from 16GB, a 125B parameter model that would need 100GB+ memory/RAM, thanks to expert-offloading/ssd-streaming. Easy to install/update, and mac-native using MLX and Swift.

It ships with auto-mode, which makes a good tradeoff between memory usage and speed.

I'll be implementing and porting the MTP module for speculative decoding next

Local models really are the future of computing!

Show HN: Slotstream, run Qwen3.8-Flash-Next 4-bit on a low-memory Mac
github.com