Earlier quoted context omitted.
Because sensible people just use the cloud at this point, you can probably get several years of training for $6000
It buys you approximately two days (with reservation discount) of a single p5.48xlarge instance, which has 2TB of RAM, and 640GB of VRAM in 8x H100 cards. In fact that is the pricing example they use: https://aws.amazon.com/ec2/capacityblocks/pricing/
Run DeepSeek R1 Dynamic 1.58-bit
251–260 of 346 posts
Re: Run DeepSeek R1 Dynamic 1.58-bit
#252Earlier quoted context omitted.
Just don’t ask it about anything related to Tiananmen square or president Pooh.. I’d guess they didn’t quite a bit of fine tuning to censor some more sensitive topics which probably impacts the output quality for other non technical subjects.
Would fine-tuning by using a LoRA paper over the censorship to a large degree?
The people making the model probably don't really give a shit about politics and just did the minimum to avoid being embarassed, but if people start jailbreaking it they will be forced to care.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#253Earlier quoted context omitted.
I ran whatever version Ollama downloaded on a 3070ti (laptop version). It's reasonably fast. Generative stuff can get weird if you do prompts like "in the style of" or "a new episode of" because it doesn't seem to have much pop culture in its training data. It knows the Stargate movie, for example, and seems to have the IMDB info for the series, but goes absolutely ham trying to summarize the series. This line in the…
It is hilariously bad at writing erotica when I've used jailbreaks on it. It's knowledge is the equivalent of a 1980s college kid with no access to pornography who watched an R rated movie once.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#254Earlier quoted context omitted.
I just ran it up on 48gb (2x 3090) + overflow into CPU RAM and it runs at around 4tk/s (only a little 8k context size though) which while absolutely not something I'd personally use daily - it is actually usable.
I have similar set-up - can you help out with running it? Was it in ollama? EDIT: It seems that original authors provided a nice write-up: https://unsloth.ai/blog/deepseekr1-dynamic#:~:text=%F0%9F%96...
Re: Run DeepSeek R1 Dynamic 1.58-bit
#255Earlier quoted context omitted.
Oh the repetition issue is only on the non dynamic quants :) If you do dynamic quantization and use the 1.58bit dynamic quantized model the repetition issue fully disappears! Min_p = 0.05 was a way I found to counteract the 1.58bit model generating singular incorrect tokens which happen around 1 token per 8000!
min_p is great, do you apply a small amount of temperate as well?
Re: Run DeepSeek R1 Dynamic 1.58-bit
#256Earlier quoted context omitted.
Laptops get stolen on a train? An enclosed, single-direction space that only occasionally allows you to exit between infrequent, long-distance stops? A thing that contains ticket inspectors and a literal guard? How many laptops have you personally seen be stolen on a train?
Yes, all the time. It's happened to two people I know, in France and in the US. People get up to use the bathroom or the cafe car, the laptop is left behind for ten minutes, one of the train stops is while they're away from their seat, and someone sees an opportunity, snags it, and gets off at the stop. This is an actual thing. And if it's worth a thousand bucks then it's very much worth getting off at an earlier sto…
My work policies state you simply cannot leave your laptop out of sight for any period unless it's in a secure location (work|home). I feel the same way for my personal laptop as well.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#257Earlier quoted context omitted.
Do you know any? But here's my advice: drop the fallacious arguments and try something more honest.
my argument isn’t fallacious - it is logical: we can learn/use evidence from something without presuming it is all knowing. you are putting words in others mouths that they did not say
Still, his implied assertion that markets that markets can often behave irrationally, and can't be used as evidence of technical matters, seems pretty valid to me.
But I suppose you could see it as a sign that something is at least temporarily "generally accepted" among investors. That doesn't mean it's generally accepted among AI researchers, though.
Although I thought it was $6M rather than $5M, and that that was only the last step, and not the total investment. What does seem to be generally accepted among investors that this isn't good news for NVidia's profits, but that still doesn't mean that all the specific facts are generally accepted.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#258Wow, an 80% reduction in size for DeepSeek-R1 is just amazing! It's fantastic to see such large models becoming more accessible to those of us who don't have access to top-tier hardware. This kind of optimization opens up so many possibilities for experimenting at home. I'm impressed by the 140 tokens per second speed with the 1.58-bit quantization running on dual H100s. That kind of performance makes the model pract…
Btw completely off topic, but your comment triggered the internal classification in my brain, and it looks like AI-generated. Not accusing you anything. Could be that you happen to write in a way similar to LLMs. Could be that we are influenced by LLM writing styles and are writing more and more like LLMs. Could be that the difference between LLM generated content and human-generated content is getting smaller and ha…
Re: Run DeepSeek R1 Dynamic 1.58-bit
#259Wow, an 80% reduction in size for DeepSeek-R1 is just amazing! It's fantastic to see such large models becoming more accessible to those of us who don't have access to top-tier hardware. This kind of optimization opens up so many possibilities for experimenting at home. I'm impressed by the 140 tokens per second speed with the 1.58-bit quantization running on dual H100s. That kind of performance makes the model pract…
Btw completely off topic, but your comment triggered the internal classification in my brain, and it looks like AI-generated. Not accusing you anything. Could be that you happen to write in a way similar to LLMs. Could be that we are influenced by LLM writing styles and are writing more and more like LLMs. Could be that the difference between LLM generated content and human-generated content is getting smaller and ha…
Amazing that OP confirmed you're correct (and good use of LLM @OP).
Re: Run DeepSeek R1 Dynamic 1.58-bit
#260Earlier quoted context omitted.
> Is the claim that it only took $5M to train generally accepted? Based on Nvidia being down 18% yesterday I would say the claim is generally accepted.
Because the markets are rational, all-knowing, and have never been wrong?