Live data from Hacker News

Run DeepSeek R1 Dynamic 1.58-bit

unsloth.ai

71–80 of 346 posts

Re: Run DeepSeek R1 Dynamic 1.58-bit

#71
post #58
post #38

Earlier quoted context omitted.

Honestly, if you have a residence of some kind and an Internet connection, you don't need to bring your beefy computer with you everywhere. It is cool to be able to have ridiculously powerful mobile computers, but I don't think I would ever be willing to take a $6,000 laptop anywhere it has a decent chance of being stolen.

Do you live in a third world country? If so I might agree, but otherwise trains are perfectly safe.

I am very happy for you, but laptops get stolen in public in most countries.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#72
post #34

As someone who is out of the loop, what’s the verdict on R1? Was anyone able to reproduce the results yet? Is the claim that it only took $5M to train generally accepted? It’s a very bold claim which is really shaking up the markets, so I can’t help but wonder if it was even verified at this point.

[deleted]

Re: Run DeepSeek R1 Dynamic 1.58-bit

#73
post #71
post #58

Earlier quoted context omitted.

Do you live in a third world country? If so I might agree, but otherwise trains are perfectly safe.

I am very happy for you, but laptops get stolen in public in most countries.

Laptops get stolen on a train? An enclosed, single-direction space that only occasionally allows you to exit between infrequent, long-distance stops? A thing that contains ticket inspectors and a literal guard?

How many laptops have you personally seen be stolen on a train?

Re: Run DeepSeek R1 Dynamic 1.58-bit

#74
post #40

Earlier quoted context omitted.

> Is the claim that it only took $5M to train generally accepted? Based on Nvidia being down 18% yesterday I would say the claim is generally accepted.

Because the markets are rational, all-knowing, and have never been wrong?

That was not the question.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#75
post #37

Earlier quoted context omitted.

So I'm thinking, inference seems mostly memory bound. With a fast CPU (for example 7950x with 16 cores), and 256GB of RAM (seems to be the max), shouldn't that give you plenty of ability to run the largest models (albeit a bit slowly). It seems that AMD Epyc CPUs support terabytes of ram, some are as cheap as 1000 EUR. why not just run the full R1 model on that - seems that it would be much cheaper than multiple of t…

The bottleneck is mainly memory bandwidth. AMD EPYC hw is appealing for local inference because it has a higher memory bandwidth than desktop gear (because 8-12 memory channels vs 2 on almost everything else), but not as fast as the Apple architectures and nowhere near VRAM speeds. If you want to drastically exceed ~3-5 tokens/s on 70b-q4 models, you usually still need GPUs.

On Zen5 you also get AVX512 which llamafile takes advantage of for drastically improved speeds during prompt processing, at least. And the 12 channel Epycs actually seem to have more memory bandwidth available than the Apple M series. Especially considering it's all available to the CPU as opposed to just some portion of it.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#76

An 80% size reduction is no joke, and the fact that the 1.58-bit version runs on dual H100s at 140 tokens/s is kind of mind-blowing. That said, I’m still skeptical about how practical this really is for most people. Like, yeah, you can run it on 24GB VRAM or even with just 20GB RAM, but "slow" is an understatement—those speeds would make even the most patient person throw their hands up. And then there’s the whole re…

> That said, I’m still skeptical about how practical this really is for most people.

I'm running Open WebUI for months now for me and some friends as a front-end to one of the API providers (deepinfra in my case, but there are many others, see https://artificialanalysis.ai/).

Having 1.58-bit is very practical for me. I'm looking much forward to the API provider adding this model to their system. They also added a Llama turbo (also quantized) a few months back so I have good hopes.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#78
post #73
post #71

Earlier quoted context omitted.

I am very happy for you, but laptops get stolen in public in most countries.

Laptops get stolen on a train? An enclosed, single-direction space that only occasionally allows you to exit between infrequent, long-distance stops? A thing that contains ticket inspectors and a literal guard? How many laptops have you personally seen be stolen on a train?

People stabbed maybe, but that tends to be more sports related than laptop related. (Yes on a national line(!))

Re: Run DeepSeek R1 Dynamic 1.58-bit

#79
post #68
post #26

Earlier quoted context omitted.

Can I use that on the train though? I can with a 128GB MacBook, without it sounding like a helicopter taking off as well.

You can use a desktop computer on a train if it's one with power outlets. Might get some funny looks, but I've seen it happen (or at least pictures). :)

Only time I've seen that done was with assistive tech and I do sympathise that those setups are difficult enough with desktops

Re: Run DeepSeek R1 Dynamic 1.58-bit

#80
post #78
post #73

Earlier quoted context omitted.

Laptops get stolen on a train? An enclosed, single-direction space that only occasionally allows you to exit between infrequent, long-distance stops? A thing that contains ticket inspectors and a literal guard? How many laptops have you personally seen be stolen on a train?

People stabbed maybe, but that tends to be more sports related than laptop related. (Yes on a national line(!))

Yeah, a long, enclosed space with no exits is more amenable to drunken violence than petty theft.
Post reply on HN