> DeepSeek-R1 has been making waves recently by rivaling OpenAI's O1 reasoning model while being fully open-source. Do we finally have a model with access to the training architecture and training data set, or are we still calling non-reproducible binary blobs without source form open-source?
Run DeepSeek R1 Dynamic 1.58-bit
121–130 of 346 posts
Re: Run DeepSeek R1 Dynamic 1.58-bit
#122For anyone wondering why "1.58" bits: 2^1.58496... = 3. The weights have one of the three states {-1, 0, 1}.
They say something else: > We managed to selectively quantize certain layers to higher bits (like 4bit), and leave most MoE layers (like those used in GPT-4) to 1.5bit
Re: Run DeepSeek R1 Dynamic 1.58-bit
#123Earlier quoted context omitted.
I am very happy for you, but laptops get stolen in public in most countries.
Laptops get stolen on a train? An enclosed, single-direction space that only occasionally allows you to exit between infrequent, long-distance stops? A thing that contains ticket inspectors and a literal guard? How many laptops have you personally seen be stolen on a train?
Obviously, nobody steals things while the train is in motion. They wait until the train is about to leave the station, snatch a phone or handbag and jump out just as the door is closing. The train leaves, the thief blends in with other passenger leaving the station, and by the time news of the theft has made it from the passengers to the driver to the station staff the thief is long gone.
Of course people drive around $6,000+ cars all the time, so....
Re: Run DeepSeek R1 Dynamic 1.58-bit
#124Random observation 1: I was running DeepSeek yesterday on my Linux with a RTX 4090 and I noticed that the models should fit into VRAM, which is 24GB. Or they are simply slow. So the Apple shared memory architecture has an advantage here. A 192GB Mx Ultra can load and process large models efficiently. Random observation 2: It's time to cancel the OpenAI subscription.
Oh yes 192GB machines should be able these quants (131GB for 1.58bit, 158GB for 1.73bit, 183GB for 2.22bit) well :)
Can you release slightly bigger quant versions? Would enjoy something that runs well on 8x32 v100 and 8x80 A100.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#125Earlier quoted context omitted.
While 192GB of ram is appealing, it's also quite expensive at $6000. For that price I rather buy a system with 5 used 3090s, which while being "only" 120GB of VRAM, you benefit from much faster tokens/s and prompt processing speed (the macs are notoriously slow at consuming large contexts).
Can I use that on the train though? I can with a 128GB MacBook, without it sounding like a helicopter taking off as well.
What kind of timescale do you expect to be able to train a useful LLM with that?
Re: Run DeepSeek R1 Dynamic 1.58-bit
#126Earlier quoted context omitted.
> They do this so you'll have to buy several cards for your AI workstation. AFAIK you can't do that with newer consumer cards, which is why this became an annoyance. Even a RTX 4070 Ti with its 12 GB would be fine, if you could easily stack a bunch of them like you used to be able with older cards.
It's "easy" if you have a place to build an open frame rig with riser cables and whatnot. I can't do that, so I'm going the single slot waterblock route, which unfortunately rules out 3090s due to the memory on the back side of the PCB. It's very frustrating.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#127Random observation 1: I was running DeepSeek yesterday on my Linux with a RTX 4090 and I noticed that the models should fit into VRAM, which is 24GB. Or they are simply slow. So the Apple shared memory architecture has an advantage here. A 192GB Mx Ultra can load and process large models efficiently. Random observation 2: It's time to cancel the OpenAI subscription.
While 192GB of ram is appealing, it's also quite expensive at $6000. For that price I rather buy a system with 5 used 3090s, which while being "only" 120GB of VRAM, you benefit from much faster tokens/s and prompt processing speed (the macs are notoriously slow at consuming large contexts).
That's because it's Apple. It time to start moving to AMD systems with shared memory. My Zen 3 APU system has 64GB these days and its a mini ITX board.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#128site is javascript walled 80%? On 2 H100 only? To get near chatgpt 4? Seriously? The 671B version??
I use Qubes OS to protect myself from the JS.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#129Earlier quoted context omitted.
Laptops get stolen on a train? An enclosed, single-direction space that only occasionally allows you to exit between infrequent, long-distance stops? A thing that contains ticket inspectors and a literal guard? How many laptops have you personally seen be stolen on a train?
Yes, all the time. It's happened to two people I know, in France and in the US. People get up to use the bathroom or the cafe car, the laptop is left behind for ten minutes, one of the train stops is while they're away from their seat, and someone sees an opportunity, snags it, and gets off at the stop. This is an actual thing. And if it's worth a thousand bucks then it's very much worth getting off at an earlier sto…
So, yes, theft on trains for people that think they are 100% safe are a thing, but applying the same idea (to assume something is 100% safe and not be cautious) I wonder how do such people use the internet...
Re: Run DeepSeek R1 Dynamic 1.58-bit
#130Earlier quoted context omitted.
While 192GB of ram is appealing, it's also quite expensive at $6000. For that price I rather buy a system with 5 used 3090s, which while being "only" 120GB of VRAM, you benefit from much faster tokens/s and prompt processing speed (the macs are notoriously slow at consuming large contexts).
>> While 192GB of ram is appealing, it's also quite expensive at $6000. That's because it's Apple. It time to start moving to AMD systems with shared memory. My Zen 3 APU system has 64GB these days and its a mini ITX board.