Live data from Hacker News

Run DeepSeek R1 Dynamic 1.58-bit

unsloth.ai

51–60 of 346 posts

Re: Run DeepSeek R1 Dynamic 1.58-bit

#51
post #40
post #34

As someone who is out of the loop, what’s the verdict on R1? Was anyone able to reproduce the results yet? Is the claim that it only took $5M to train generally accepted? It’s a very bold claim which is really shaking up the markets, so I can’t help but wonder if it was even verified at this point.

> Is the claim that it only took $5M to train generally accepted? Based on Nvidia being down 18% yesterday I would say the claim is generally accepted.

> Nvidia being down 18%

The only part of DeepSeek-R1 I do not like. I hope it's over, but I am not holding my breath.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#52
It is going to be truly fucking revolutionary if open-source models are and continue to be able to challenge the state of the art. My big philosophical concern is that AI locks Capital into an absolutely supreme and insurmountable lead over Labour, and into the hands of oligarchs, and the possibility of a future where that's not case feels amazing. It pleases me greatly that this has Trump riled up too, because I think it means he's much less likely to allow existing US model-makers to build moats, as I think he's -- even as a man who I don't think believes in very much -- absolutely unwilling to let the Chinese get the drop on him over this.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#53

An 80% size reduction is no joke, and the fact that the 1.58-bit version runs on dual H100s at 140 tokens/s is kind of mind-blowing. That said, I’m still skeptical about how practical this really is for most people. Like, yeah, you can run it on 24GB VRAM or even with just 20GB RAM, but "slow" is an understatement—those speeds would make even the most patient person throw their hands up. And then there’s the whole re…

Not everyone needs the largest model. There are variations or R1 with fewer parameters that can easily run on consumer hardware. With 80% size reduction you could run 70B on 8-bit on an RTX 3090.

Other than that, if you really need the big one you can get six 3090s and you're good to go. It's not cheap, but you're running a ChatGPT equivalent model from your basement. A year ago this was a wetdream for most enthusiasts.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#54

Random observation 1: I was running DeepSeek yesterday on my Linux with a RTX 4090 and I noticed that the models should fit into VRAM, which is 24GB. Or they are simply slow. So the Apple shared memory architecture has an advantage here. A 192GB Mx Ultra can load and process large models efficiently. Random observation 2: It's time to cancel the OpenAI subscription.

I disagree with cancelling the OpenAI subscription. I've been getting some help from o1 for both python and php recently, and o1 was doing massively better for the python stuff (it ran, deepseeks didn't and wont with prompt refinement).

Were you running a local model?

Re: Run DeepSeek R1 Dynamic 1.58-bit

#55
post #34

As someone who is out of the loop, what’s the verdict on R1? Was anyone able to reproduce the results yet? Is the claim that it only took $5M to train generally accepted? It’s a very bold claim which is really shaking up the markets, so I can’t help but wonder if it was even verified at this point.

Huggingface is working on reproducing it: https://github.com/huggingface/open-r1

Re: Run DeepSeek R1 Dynamic 1.58-bit

#56
post #40
post #34

As someone who is out of the loop, what’s the verdict on R1? Was anyone able to reproduce the results yet? Is the claim that it only took $5M to train generally accepted? It’s a very bold claim which is really shaking up the markets, so I can’t help but wonder if it was even verified at this point.

> Is the claim that it only took $5M to train generally accepted? Based on Nvidia being down 18% yesterday I would say the claim is generally accepted.

It is still unconfirmed since no one outside of deepseek reproduced it.

If confirmed, Nvidia could go down even more

Re: Run DeepSeek R1 Dynamic 1.58-bit

#57
post #39

Wow, an 80% reduction in size for DeepSeek-R1 is just amazing! It's fantastic to see such large models becoming more accessible to those of us who don't have access to top-tier hardware. This kind of optimization opens up so many possibilities for experimenting at home. I'm impressed by the 140 tokens per second speed with the 1.58-bit quantization running on dual H100s. That kind of performance makes the model pract…

I was pleasantly surprised by 140 tokens/s as well! I literally thought I did something wrong but it was real!

Re: Run DeepSeek R1 Dynamic 1.58-bit

#58
post #38
post #26

Earlier quoted context omitted.

Can I use that on the train though? I can with a 128GB MacBook, without it sounding like a helicopter taking off as well.

Honestly, if you have a residence of some kind and an Internet connection, you don't need to bring your beefy computer with you everywhere. It is cool to be able to have ridiculously powerful mobile computers, but I don't think I would ever be willing to take a $6,000 laptop anywhere it has a decent chance of being stolen.

Do you live in a third world country? If so I might agree, but otherwise trains are perfectly safe.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#59

Earlier quoted context omitted.

5x 3090 is also much more power hungry?

For personal usage, does it matter though? In most places residential electricity is cheap compared to everything else. In a DC context I feel it matters a lot more compared to the capex.

1x 3090 (350W power limit) already makes it feel like I'm running a fan heater under my desk, 5x would be nuts.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#60
post #40
post #34

As someone who is out of the loop, what’s the verdict on R1? Was anyone able to reproduce the results yet? Is the claim that it only took $5M to train generally accepted? It’s a very bold claim which is really shaking up the markets, so I can’t help but wonder if it was even verified at this point.

> Is the claim that it only took $5M to train generally accepted? Based on Nvidia being down 18% yesterday I would say the claim is generally accepted.

Because the markets are rational, all-knowing, and have never been wrong?
Post reply on HN