Live data from Hacker News

Run DeepSeek R1 Dynamic 1.58-bit

unsloth.ai

111–120 of 346 posts

Re: Run DeepSeek R1 Dynamic 1.58-bit

#111
post #63
post #53

Earlier quoted context omitted.

Not everyone needs the largest model. There are variations or R1 with fewer parameters that can easily run on consumer hardware. With 80% size reduction you could run 70B on 8-bit on an RTX 3090. Other than that, if you really need the big one you can get six 3090s and you're good to go. It's not cheap, but you're running a ChatGPT equivalent model from your basement. A year ago this was a wetdream for most enthusias…

I ran whatever version Ollama downloaded on a 3070ti (laptop version). It's reasonably fast. Generative stuff can get weird if you do prompts like "in the style of" or "a new episode of" because it doesn't seem to have much pop culture in its training data. It knows the Stargate movie, for example, and seems to have the IMDB info for the series, but goes absolutely ham trying to summarize the series. This line in the…

It is hilariously bad at writing erotica when I've used jailbreaks on it. It's knowledge is the equivalent of a 1980s college kid with no access to pornography who watched an R rated movie once.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#113
Danielhanchen, your work is continually impressive. Unsloth is great, and I’m repeatedly amazed at your ability to get up to speed on a new model within hours of its release, and often fix bugs in the default implementation. At this point, I think serious labs should give you a few hour head start just to iron out their kinks!

Re: Run DeepSeek R1 Dynamic 1.58-bit

#114
post #88
post #73

Earlier quoted context omitted.

Laptops get stolen on a train? An enclosed, single-direction space that only occasionally allows you to exit between infrequent, long-distance stops? A thing that contains ticket inspectors and a literal guard? How many laptops have you personally seen be stolen on a train?

Only on Hacker News would I have someone arguing with me that laptop theft is not a concern. You know what, you win. It's your $6,000 laptop, not mine.

Yes! This goes in my forthcoming blog post "Only on Hacker News..."

Yesterday's entry: "... kind of a mind flex that you noted you used Meta Stories glasses to take that photo."

Re: Run DeepSeek R1 Dynamic 1.58-bit

#115

An 80% size reduction is no joke, and the fact that the 1.58-bit version runs on dual H100s at 140 tokens/s is kind of mind-blowing. That said, I’m still skeptical about how practical this really is for most people. Like, yeah, you can run it on 24GB VRAM or even with just 20GB RAM, but "slow" is an understatement—those speeds would make even the most patient person throw their hands up. And then there’s the whole re…

At my work, we self-host some models and have found that for anything remotely similar to RAG or use cases that are very specific, the quantized models have proven to be more than sufficient. This helps us keep them running on smaller infra and generally lower costs

Personally I've noticed major changes in performance between different quantisations of the same model.

Mistral's large 123B model works well (but slowly) at 4-bit quantisation, but if I knock it down to 2.5-bit quantisation for speed, performance drops to the point where I'm better off with a 70B 4-bit model.

This makes me reluctant to evaluate new models in heavily quantised forms, as you're measuring the quantisation more than the actual model.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#116
post #56
post #40

Earlier quoted context omitted.

> Is the claim that it only took $5M to train generally accepted? Based on Nvidia being down 18% yesterday I would say the claim is generally accepted.

It is still unconfirmed since no one outside of deepseek reproduced it. If confirmed, Nvidia could go down even more

based on information and background they thoroughly gave when releasing their research its pretty easy to put together that it did take them significantly less resources to train this model. only having specific parameters available at a time instead of activating everything all at once is pretty ingenious.

that and they just happened to be undergoing a large scale "cyber attack"

Re: Run DeepSeek R1 Dynamic 1.58-bit

#118
post #65

Earlier quoted context omitted.

The real insult here is graphics card vendors refusing to make ones with more than 24GB for several years now. They do this so you'll have to buy several cards for your AI workstation. Hopefully Apple eating their lunch fixes this.

> They do this so you'll have to buy several cards for your AI workstation. AFAIK you can't do that with newer consumer cards, which is why this became an annoyance. Even a RTX 4070 Ti with its 12 GB would be fine, if you could easily stack a bunch of them like you used to be able with older cards.

It's "easy" if you have a place to build an open frame rig with riser cables and whatnot. I can't do that, so I'm going the single slot waterblock route, which unfortunately rules out 3090s due to the memory on the back side of the PCB. It's very frustrating.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#119
post #21

Random observation 1: I was running DeepSeek yesterday on my Linux with a RTX 4090 and I noticed that the models should fit into VRAM, which is 24GB. Or they are simply slow. So the Apple shared memory architecture has an advantage here. A 192GB Mx Ultra can load and process large models efficiently. Random observation 2: It's time to cancel the OpenAI subscription.

While 192GB of ram is appealing, it's also quite expensive at $6000. For that price I rather buy a system with 5 used 3090s, which while being "only" 120GB of VRAM, you benefit from much faster tokens/s and prompt processing speed (the macs are notoriously slow at consuming large contexts).

I think just getting nvidia Project Digits might be the best option. A lot of people when it was announced were underwhelmed. But I think now it could be just the thing for people making their own ai home servers.

https://www.nvidia.com/en-us/project-digits/

Re: Run DeepSeek R1 Dynamic 1.58-bit

#120

Earlier quoted context omitted.

You mean a tight, enclosed, single-direction space, crowded with people who are tired, and/or trying to relax, and/or thinking about the destination, and/or otherwise not particularly focused after hours of travel; a thing that contains ticket inspectors that show up every now and then to check tickets , and from which passengers embark and disembark at dozens point along the length of the thing, simultaneously, with…

Pickpocketing is a very different proposition. They relying on a lack of awareness, taking your wallet and being long gone before you’ve even noticed. If someone steals your laptop from in front of you without you even noticing I’d suggest that one is on you. FWIW I’ve used my laptop on the train plenty, I’ve never had anything stolen nor felt in any danger of it.

But would you consider leaving it unattended on your seat and going for lunch to the restaurant car, or for an extended toilet break?
Post reply on HN