Live data from Hacker News

Run DeepSeek R1 Dynamic 1.58-bit

unsloth.ai

251–260 of 346 posts

Re: Run DeepSeek R1 Dynamic 1.58-bit

#251
post #155

Earlier quoted context omitted.

Because sensible people just use the cloud at this point, you can probably get several years of training for $6000

It buys you approximately two days (with reservation discount) of a single p5.48xlarge instance, which has 2TB of RAM, and 640GB of VRAM in 8x H100 cards. In fact that is the pricing example they use: https://aws.amazon.com/ec2/capacityblocks/pricing/

MI300X (RunPod) 192gb ram Hourly Rate: $2.49/hr. Break-even Point: You can rent for 2,410 hours (~100 days of non-stop-continuous use) before reaching the cost of the $6000 Mac. Mac's top out at 192GB not 2TB ;) Consideration: If your AI training requires sporadic use (e.g., a few hours daily or weekly), renting is significantly cheaper. MI300X will also get you result many times faster too, so you could probably multiply that 100 days!

Re: Run DeepSeek R1 Dynamic 1.58-bit

#252

Earlier quoted context omitted.

Just don’t ask it about anything related to Tiananmen square or president Pooh.. I’d guess they didn’t quite a bit of fine tuning to censor some more sensitive topics which probably impacts the output quality for other non technical subjects.

Would fine-tuning by using a LoRA paper over the censorship to a large degree?

Why even bother decensoring it (except academic curiosity ig)? There are a million other ways you can learn about those subjects.

The people making the model probably don't really give a shit about politics and just did the minimum to avoid being embarassed, but if people start jailbreaking it they will be forced to care.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#253
post #63

Earlier quoted context omitted.

I ran whatever version Ollama downloaded on a 3070ti (laptop version). It's reasonably fast. Generative stuff can get weird if you do prompts like "in the style of" or "a new episode of" because it doesn't seem to have much pop culture in its training data. It knows the Stargate movie, for example, and seems to have the IMDB info for the series, but goes absolutely ham trying to summarize the series. This line in the…

It is hilariously bad at writing erotica when I've used jailbreaks on it. It's knowledge is the equivalent of a 1980s college kid with no access to pornography who watched an R rated movie once.

That's like trying to assemble an Ikea bookshelf with a bulldozer. All that extra power is doing nothing for the task you're asking of it, and there are plenty of lightweight alternatives.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#254
post #67

Earlier quoted context omitted.

I just ran it up on 48gb (2x 3090) + overflow into CPU RAM and it runs at around 4tk/s (only a little 8k context size though) which while absolutely not something I'd personally use daily - it is actually usable.

I have similar set-up - can you help out with running it? Was it in ollama? EDIT: It seems that original authors provided a nice write-up: https://unsloth.ai/blog/deepseekr1-dynamic#:~:text=%F0%9F%96...

Yep that's pretty much what I did, their calculation for the layers was slightly off though, I found I could offload an extra 1-2 layers to the GPUs

Re: Run DeepSeek R1 Dynamic 1.58-bit

#255
post #69

Earlier quoted context omitted.

Oh the repetition issue is only on the non dynamic quants :) If you do dynamic quantization and use the 1.58bit dynamic quantized model the repetition issue fully disappears! Min_p = 0.05 was a way I found to counteract the 1.58bit model generating singular incorrect tokens which happen around 1 token per 8000!

min_p is great, do you apply a small amount of temperate as well?

The recommended temperature from DeepSeek is 0.6 so I leave it at that!

Re: Run DeepSeek R1 Dynamic 1.58-bit

#256
post #73

Earlier quoted context omitted.

Laptops get stolen on a train? An enclosed, single-direction space that only occasionally allows you to exit between infrequent, long-distance stops? A thing that contains ticket inspectors and a literal guard? How many laptops have you personally seen be stolen on a train?

Yes, all the time. It's happened to two people I know, in France and in the US. People get up to use the bathroom or the cafe car, the laptop is left behind for ten minutes, one of the train stops is while they're away from their seat, and someone sees an opportunity, snags it, and gets off at the stop. This is an actual thing. And if it's worth a thousand bucks then it's very much worth getting off at an earlier sto…

Ok - that's really poor opsec. If I'm going to the bathroom in a train with my laptop (whether it's expensive or not - it has access to all my stuff - which is arguably more valuable), I'll sleep it, put it in my backpack and take the backpack to the bathroom with me.

My work policies state you simply cannot leave your laptop out of sight for any period unless it's in a secure location (work|home). I feel the same way for my personal laptop as well.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#257
post #220

Earlier quoted context omitted.

Do you know any? But here's my advice: drop the fallacious arguments and try something more honest.

my argument isn’t fallacious - it is logical: we can learn/use evidence from something without presuming it is all knowing. you are putting words in others mouths that they did not say

I'm sorry, I thought you introduced the "all-knowing" out of nowhere, but this was indeed mentioned by willsmith72. I'd missed that.

Still, his implied assertion that markets that markets can often behave irrationally, and can't be used as evidence of technical matters, seems pretty valid to me.

But I suppose you could see it as a sign that something is at least temporarily "generally accepted" among investors. That doesn't mean it's generally accepted among AI researchers, though.

Although I thought it was $6M rather than $5M, and that that was only the last step, and not the total investment. What does seem to be generally accepted among investors that this isn't good news for NVidia's profits, but that still doesn't mean that all the specific facts are generally accepted.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#258
post #39

Wow, an 80% reduction in size for DeepSeek-R1 is just amazing! It's fantastic to see such large models becoming more accessible to those of us who don't have access to top-tier hardware. This kind of optimization opens up so many possibilities for experimenting at home. I'm impressed by the 140 tokens per second speed with the 1.58-bit quantization running on dual H100s. That kind of performance makes the model pract…

Btw completely off topic, but your comment triggered the internal classification in my brain, and it looks like AI-generated. Not accusing you anything. Could be that you happen to write in a way similar to LLMs. Could be that we are influenced by LLM writing styles and are writing more and more like LLMs. Could be that the difference between LLM generated content and human-generated content is getting smaller and ha…

[deleted]

Re: Run DeepSeek R1 Dynamic 1.58-bit

#259
post #39

Wow, an 80% reduction in size for DeepSeek-R1 is just amazing! It's fantastic to see such large models becoming more accessible to those of us who don't have access to top-tier hardware. This kind of optimization opens up so many possibilities for experimenting at home. I'm impressed by the 140 tokens per second speed with the 1.58-bit quantization running on dual H100s. That kind of performance makes the model pract…

Btw completely off topic, but your comment triggered the internal classification in my brain, and it looks like AI-generated. Not accusing you anything. Could be that you happen to write in a way similar to LLMs. Could be that we are influenced by LLM writing styles and are writing more and more like LLMs. Could be that the difference between LLM generated content and human-generated content is getting smaller and ha…

Very funny, I didn't mentally jump to LLM, but the language was so lifeless that I stopped reading.

Amazing that OP confirmed you're correct (and good use of LLM @OP).

Re: Run DeepSeek R1 Dynamic 1.58-bit

#260
post #40

Earlier quoted context omitted.

> Is the claim that it only took $5M to train generally accepted? Based on Nvidia being down 18% yesterday I would say the claim is generally accepted.

Because the markets are rational, all-knowing, and have never been wrong?

as opposed to HN comments??
Post reply on HN