Run DeepSeek R1 Dynamic 1.58-bit
231–240 of 346 posts
Re: Run DeepSeek R1 Dynamic 1.58-bit
#232Re: Run DeepSeek R1 Dynamic 1.58-bit
#233Earlier quoted context omitted.
> Is the claim that it only took $5M to train generally accepted? Based on Nvidia being down 18% yesterday I would say the claim is generally accepted.
Because the markets are rational, all-knowing, and have never been wrong?
It may provide a financial opportunity for someone who disagrees with that aggregated opinion though.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#234An 80% size reduction is no joke, and the fact that the 1.58-bit version runs on dual H100s at 140 tokens/s is kind of mind-blowing. That said, I’m still skeptical about how practical this really is for most people. Like, yeah, you can run it on 24GB VRAM or even with just 20GB RAM, but "slow" is an understatement—those speeds would make even the most patient person throw their hands up. And then there’s the whole re…
Re: Run DeepSeek R1 Dynamic 1.58-bit
#235An 80% size reduction is no joke, and the fact that the 1.58-bit version runs on dual H100s at 140 tokens/s is kind of mind-blowing. That said, I’m still skeptical about how practical this really is for most people. Like, yeah, you can run it on 24GB VRAM or even with just 20GB RAM, but "slow" is an understatement—those speeds would make even the most patient person throw their hands up. And then there’s the whole re…
>Like, I get that shared memory architectures like a 192GB Mac Ultra are a big deal, but who’s dropping $6,000+ on that setup? AMD strix halo APU will have quad channel memory and will launch soon so expect these kinds of setups available for much less. Apple is charging an arm and a leg for memory upgrades, hopefully we get competition soon. From what I saw at CES OEMs are paying attention to this use case as well -…
Here's hoping the Nvidia Digit (GB10 chip) has a 512 bit or 1024 bit wide interface, otherwise the Strix Halo will be the best you can do if you don't get the Mac Ultra.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#236I love the original DeepSeek model, but the distilled versions are too dumb usually. I'm excited to try my own queries on it.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#237Hi small comment, please remember in china many things are sponsored by or subsidized by the government. "We[china] can do it for less.." , "it's cheaper in china.." only means the government gave us a pile of cash and help to get here . I 100% expect some downvotes from the ccp.
Always happy to oblige when someone insinuates that any critics must be government agents
Re: Run DeepSeek R1 Dynamic 1.58-bit
#238Earlier quoted context omitted.
You want to bet? The panic around deepseek is getting completely disconnected from reality. Don’t get me wrong what DS did is great, but anyone thinking this reshape the fundamental trend of scaling laws and make compute irrelevant is dead wrong. I’m sure OpenAI doesn’t really enjoy the PR right now, but guess what OpenAI/Google/Meta/Anthropic can do if you give them a recipe for 11x more efficient training ? They ca…
> You want to bet? Why would anyone bet? They can just short the OpenAI / MS stocks, and see in a few months if they were right or not.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#239Earlier quoted context omitted.
You want to bet? The panic around deepseek is getting completely disconnected from reality. Don’t get me wrong what DS did is great, but anyone thinking this reshape the fundamental trend of scaling laws and make compute irrelevant is dead wrong. I’m sure OpenAI doesn’t really enjoy the PR right now, but guess what OpenAI/Google/Meta/Anthropic can do if you give them a recipe for 11x more efficient training ? They ca…
> You want to bet? Why would anyone bet? They can just short the OpenAI / MS stocks, and see in a few months if they were right or not.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#240Random observation 1: I was running DeepSeek yesterday on my Linux with a RTX 4090 and I noticed that the models should fit into VRAM, which is 24GB. Or they are simply slow. So the Apple shared memory architecture has an advantage here. A 192GB Mx Ultra can load and process large models efficiently. Random observation 2: It's time to cancel the OpenAI subscription.
The real insult here is graphics card vendors refusing to make ones with more than 24GB for several years now. They do this so you'll have to buy several cards for your AI workstation. Hopefully Apple eating their lunch fixes this.