Run DeepSeek R1 Dynamic 1.58-bit
unsloth.ai
Run DeepSeek R1 Dynamic 1.58-bit
1–10 of 346 posts
Re: Run DeepSeek R1 Dynamic 1.58-bit
#2Re: Run DeepSeek R1 Dynamic 1.58-bit
#3For example, I imagine a strong MoE base with 16 billion active parameters and 6 or 7 experts would keep a good performance while being possible to run on 128GB RAM macbooks.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#4This is really interesting insight (although other works cover this as well). I am particularly amused by the process by which the authors of this blog post arrived at these particular seeds. Good work nonetheless!
Re: Run DeepSeek R1 Dynamic 1.58-bit
#5Re: Run DeepSeek R1 Dynamic 1.58-bit
#6Re: Run DeepSeek R1 Dynamic 1.58-bit
#7Would be great if the next generation of base models was designed to be inferred with 128GB of VRAM while 8bit quantized (which would fit in the consumer hardware class). For example, I imagine a strong MoE base with 16 billion active parameters and 6 or 7 experts would keep a good performance while being possible to run on 128GB RAM macbooks.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#8In general, how do you run these big models on cloud hardware? Do you cut them up layer-wise and run slices of layers on individual A100/H100s?
Re: Run DeepSeek R1 Dynamic 1.58-bit
#9Would be great if the next generation of base models was designed to be inferred with 128GB of VRAM while 8bit quantized (which would fit in the consumer hardware class). For example, I imagine a strong MoE base with 16 billion active parameters and 6 or 7 experts would keep a good performance while being possible to run on 128GB RAM macbooks.
Would be great, but unfortunately i think intelligence at that compute scale will be limit by hardware not its model. Though at hardware limit I would expect it to be roughly human level especially if optimized for a particular domain.
Maybe using a strong reasoning model such as R1 the next generation, even more performance can be extracted from smaller models.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#10In general, how do you run these big models on cloud hardware? Do you cut them up layer-wise and run slices of layers on individual A100/H100s?