Live data from Hacker News

Run DeepSeek R1 Dynamic 1.58-bit

unsloth.ai

321–330 of 346 posts

Re: Run DeepSeek R1 Dynamic 1.58-bit

#321
post #295

Earlier quoted context omitted.

I have MLCCHAT on my old Note 9 phone. It is actually still a great phone, but has 5GB RAM. Running an on device model is the first and only use case the RAM actually matters. And it has a headphone jack, OK? I just hate Bluetooth earbuds. And yeah, it isna problem, but I digress. When I run a 2.5B model, I get respectable output. Takes a minute or two to process the context, then output begins at somewhere on the or…

> First aid, how to make fires, materials and uses This scares me more than it should... Please do not trust an AI in actual life and death situations... Sure if it is literally your only option, but this implies you have a device on you that could make a phone call to an emergency number where a real human with real training and actually correct knowledge can assist you. Even as an avid hiker the amount of times I'v…

Of course! I do the same. However, I won't deny being able to get some information, even if I must validate it with care, jn a pinch is a great thing.

It just a tool in the tool box. Like any tool, one must respect and use it with care.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#322
post #266

Earlier quoted context omitted.

You want to bet? The panic around deepseek is getting completely disconnected from reality. Don’t get me wrong what DS did is great, but anyone thinking this reshape the fundamental trend of scaling laws and make compute irrelevant is dead wrong. I’m sure OpenAI doesn’t really enjoy the PR right now, but guess what OpenAI/Google/Meta/Anthropic can do if you give them a recipe for 11x more efficient training ? They ca…

Deepseek is not the only reason. I cancelled my OpenAI subscription because I've replaced it wholesale with Anthropic.

I replaced that with kagi, unliminted access to multiple models including Claude, O1 and V3/R1 + you also get Kagi, which was already a good deal

Re: Run DeepSeek R1 Dynamic 1.58-bit

#323
post #268

Earlier quoted context omitted.

1. You can get all the models by buying Kagi subscription (excluding o1). Includes DeepSeek models. You can also feed the assistant with search data that you can filter. 2. If you have GitHub Copilot, you get o1 chat also there. I haven't seen much value with OpenAI subscription for ages.

I have Kagi Ultimate and it is nice for this. But a cheaper suggestion would be to use OpenRouter and then use these models via Fireworks or TogetherAI. It also integrates into much more applications. AFAIK Kagi doesn't document a user facing API for the assistant feature.

Unfortunately those are both 10-15x the cost of deepseek direct.

Deepinfra is pretty cheap though as a deepseek provider.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#324
post #170

Earlier quoted context omitted.

> Hi small comment, please remember in china many things are sponsored by or subsidized by the government. "We[china] can do it for less.." , "it's cheaper in china.." only means the government gave us a pile of cash and help to get here . And that's a really important strategic advantage China has versus America, which has such an insane fixation on pure(ish) free markets and free trade that it gives away its advant…

> And that's a really important strategic advantage China has versus America, which has such an insane fixation on pure(ish) free markets and free trade that it gives away its advantages in strategic industry after strategic industry. > Some people falsely infer from the experience with the Soviet Union that freer markets always win geopolitical competition, but that's false. The data we have is 500 years of free mar…

> The data we have is 500 years of free markets in the western world and the verdict is overwhelmingly: Yes, more freedom means more winning.

No, more freedom means more winning to a point. Past that point it does not, and I'd argue that's where the US is.

> Just invite some incompetent bureaucrat over your house to dictate how you should cook and you'll quickly agree.

That's supposed to be convincing, somehow? Just invite some "competent" capitalist over to your house, and he'll sell your fishing rod in exchange for a short-term discount on fish at the supermarket, and see how well you win.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#325
post #91

Earlier quoted context omitted.

I canceled my OpenAI subscription last night, as did many many others. There were some threads in reddit with everyone chiming in they all just canceled too. imo OpenAI is done, and will go through massive cuts and probably acquired by the end of the year for a very tiny fraction of its current value.

You want to bet? The panic around deepseek is getting completely disconnected from reality. Don’t get me wrong what DS did is great, but anyone thinking this reshape the fundamental trend of scaling laws and make compute irrelevant is dead wrong. I’m sure OpenAI doesn’t really enjoy the PR right now, but guess what OpenAI/Google/Meta/Anthropic can do if you give them a recipe for 11x more efficient training ? They ca…

OpenAI issue might be that it is extremely inefficient with money (high salaries, high compute costs, high expenses, etc..). This is fine when you have an absolute monopoly as investors will throw money your way (open ai is burning cash) but once an alternative is clear, you can no longer do that.

OpenAI doesn't have an advantage in compute more than Google, Microsoft or someone with a few billions of $$.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#326

Earlier quoted context omitted.

I have Kagi Ultimate and it is nice for this. But a cheaper suggestion would be to use OpenRouter and then use these models via Fireworks or TogetherAI. It also integrates into much more applications. AFAIK Kagi doesn't document a user facing API for the assistant feature.

Unfortunately those are both 10-15x the cost of deepseek direct. Deepinfra is pretty cheap though as a deepseek provider.

That's because DeepSeek is subsidizing their API massively to get more training data.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#327
post #277

Earlier quoted context omitted.

How are you arriving at those numbers? ryzen 5500 + 7x3060 + cooling ~= 1.6 kW off the wall, at 360 GB/s memory bandwidth, and considering your lane budget, most of it will be wasted in single PCIe lanes. After-market unit price of 3060's is 200 eur, so 1600 is not good-faith cost estimate. From the looks of it, your setup is neither low-power, nor low-cost. You'd be better served with a refurbished mac studio (2022)…

The issue is that you are taking max GPU power draw, as a given. Running a LLM does not tax a GPU the same way a game does. There is a rather know Youtuber, that ran LLMs on a 4090, and the actual power draw was only 130W on the GPU. Now add that this guy has 7x3060 = 100% miner. So you know that he is running a optimized profile (underclocked). Fyi, my gaming 6800 draws 230W, but with a bit of undervolting and sacri…

"Undervolting" is a thing for 3090s where they get them down from 350 to 300W at 5% perf drop but for your case it's irrelevant because your lane budget is far too little!

> know Youtuber, that ran LLMs on a 4090, and the actual power draw was only 130W on the GPU.

Well, let's see his video. He must be using some really inefficient backend implementation if the GPU wasn't utilised like that.

I'm not running e-waste. My cards are L40S and even in basic inference, no batching with ggml cuda kernels they get to 70% util immediately.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#328
post #70
post #37

Earlier quoted context omitted.

So I'm thinking, inference seems mostly memory bound. With a fast CPU (for example 7950x with 16 cores), and 256GB of RAM (seems to be the max), shouldn't that give you plenty of ability to run the largest models (albeit a bit slowly). It seems that AMD Epyc CPUs support terabytes of ram, some are as cheap as 1000 EUR. why not just run the full R1 model on that - seems that it would be much cheaper than multiple of t…

FWIW Threadrippers go up to 1TB and Threadripper Pro up to 2TB. That's even in the lowest model of each series. (I know this because it happens to be the chip I have. Not saying you shouldn't go for Epyc if it works out better.)

Have you tried running the full R1 model with that? People in sibling comments mention high end EPYCs gor a 10K machine, but I’m curious whether it’s possible to make a 1-2K machine that could still run those big models simply because they fit in RAM.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#329
post #328
post #70

Earlier quoted context omitted.

FWIW Threadrippers go up to 1TB and Threadripper Pro up to 2TB. That's even in the lowest model of each series. (I know this because it happens to be the chip I have. Not saying you shouldn't go for Epyc if it works out better.)

Have you tried running the full R1 model with that? People in sibling comments mention high end EPYCs gor a 10K machine, but I’m curious whether it’s possible to make a 1-2K machine that could still run those big models simply because they fit in RAM.

I spent about $3000 on my machine, have the cheapest Threadripper CPU and 256GB of RAM, so no, 600GB won't fit in RAM on a $2K machine.

But everyone is using the distilled models which are much smaller.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#330
post #266

Earlier quoted context omitted.

Deepseek is not the only reason. I cancelled my OpenAI subscription because I've replaced it wholesale with Anthropic.

I replaced that with kagi, unliminted access to multiple models including Claude, O1 and V3/R1 + you also get Kagi, which was already a good deal

oh wow. I have been using kagi premium for months, and never noticed, that their AI assistant now has all the good AIs too. I was using kagi exclusively for search, and perplexity for ai stuff. I guess I can cut down on my subscriptions too. Thanks for your hint. (Also I noticed that kagi has a pwa for their ai assistent, which is also cool)
Post reply on HN