Earlier quoted context omitted.
Not everyone needs the largest model. There are variations or R1 with fewer parameters that can easily run on consumer hardware. With 80% size reduction you could run 70B on 8-bit on an RTX 3090. Other than that, if you really need the big one you can get six 3090s and you're good to go. It's not cheap, but you're running a ChatGPT equivalent model from your basement. A year ago this was a wetdream for most enthusias…
I ran whatever version Ollama downloaded on a 3070ti (laptop version). It's reasonably fast. Generative stuff can get weird if you do prompts like "in the style of" or "a new episode of" because it doesn't seem to have much pop culture in its training data. It knows the Stargate movie, for example, and seems to have the IMDB info for the series, but goes absolutely ham trying to summarize the series. This line in the…
Run DeepSeek R1 Dynamic 1.58-bit
111–120 of 346 posts
Re: Run DeepSeek R1 Dynamic 1.58-bit
#112Re: Run DeepSeek R1 Dynamic 1.58-bit
#113Re: Run DeepSeek R1 Dynamic 1.58-bit
#114Earlier quoted context omitted.
Laptops get stolen on a train? An enclosed, single-direction space that only occasionally allows you to exit between infrequent, long-distance stops? A thing that contains ticket inspectors and a literal guard? How many laptops have you personally seen be stolen on a train?
Only on Hacker News would I have someone arguing with me that laptop theft is not a concern. You know what, you win. It's your $6,000 laptop, not mine.
Yesterday's entry: "... kind of a mind flex that you noted you used Meta Stories glasses to take that photo."
Re: Run DeepSeek R1 Dynamic 1.58-bit
#115An 80% size reduction is no joke, and the fact that the 1.58-bit version runs on dual H100s at 140 tokens/s is kind of mind-blowing. That said, I’m still skeptical about how practical this really is for most people. Like, yeah, you can run it on 24GB VRAM or even with just 20GB RAM, but "slow" is an understatement—those speeds would make even the most patient person throw their hands up. And then there’s the whole re…
At my work, we self-host some models and have found that for anything remotely similar to RAG or use cases that are very specific, the quantized models have proven to be more than sufficient. This helps us keep them running on smaller infra and generally lower costs
Mistral's large 123B model works well (but slowly) at 4-bit quantisation, but if I knock it down to 2.5-bit quantisation for speed, performance drops to the point where I'm better off with a 70B 4-bit model.
This makes me reluctant to evaluate new models in heavily quantised forms, as you're measuring the quantisation more than the actual model.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#116Earlier quoted context omitted.
> Is the claim that it only took $5M to train generally accepted? Based on Nvidia being down 18% yesterday I would say the claim is generally accepted.
It is still unconfirmed since no one outside of deepseek reproduced it. If confirmed, Nvidia could go down even more
that and they just happened to be undergoing a large scale "cyber attack"
Re: Run DeepSeek R1 Dynamic 1.58-bit
#117How can you have a bit and a half exactly? It doesn't make sense.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#118Earlier quoted context omitted.
The real insult here is graphics card vendors refusing to make ones with more than 24GB for several years now. They do this so you'll have to buy several cards for your AI workstation. Hopefully Apple eating their lunch fixes this.
> They do this so you'll have to buy several cards for your AI workstation. AFAIK you can't do that with newer consumer cards, which is why this became an annoyance. Even a RTX 4070 Ti with its 12 GB would be fine, if you could easily stack a bunch of them like you used to be able with older cards.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#119Random observation 1: I was running DeepSeek yesterday on my Linux with a RTX 4090 and I noticed that the models should fit into VRAM, which is 24GB. Or they are simply slow. So the Apple shared memory architecture has an advantage here. A 192GB Mx Ultra can load and process large models efficiently. Random observation 2: It's time to cancel the OpenAI subscription.
While 192GB of ram is appealing, it's also quite expensive at $6000. For that price I rather buy a system with 5 used 3090s, which while being "only" 120GB of VRAM, you benefit from much faster tokens/s and prompt processing speed (the macs are notoriously slow at consuming large contexts).
Re: Run DeepSeek R1 Dynamic 1.58-bit
#120Earlier quoted context omitted.
You mean a tight, enclosed, single-direction space, crowded with people who are tired, and/or trying to relax, and/or thinking about the destination, and/or otherwise not particularly focused after hours of travel; a thing that contains ticket inspectors that show up every now and then to check tickets , and from which passengers embark and disembark at dozens point along the length of the thing, simultaneously, with…
Pickpocketing is a very different proposition. They relying on a lack of awareness, taking your wallet and being long gone before you’ve even noticed. If someone steals your laptop from in front of you without you even noticing I’d suggest that one is on you. FWIW I’ve used my laptop on the train plenty, I’ve never had anything stolen nor felt in any danger of it.