Live data from Hacker News

Run DeepSeek R1 Dynamic 1.58-bit

unsloth.ai

191–200 of 346 posts

Re: Run DeepSeek R1 Dynamic 1.58-bit

#191
post #91

Earlier quoted context omitted.

I canceled my OpenAI subscription last night, as did many many others. There were some threads in reddit with everyone chiming in they all just canceled too. imo OpenAI is done, and will go through massive cuts and probably acquired by the end of the year for a very tiny fraction of its current value.

You want to bet? The panic around deepseek is getting completely disconnected from reality. Don’t get me wrong what DS did is great, but anyone thinking this reshape the fundamental trend of scaling laws and make compute irrelevant is dead wrong. I’m sure OpenAI doesn’t really enjoy the PR right now, but guess what OpenAI/Google/Meta/Anthropic can do if you give them a recipe for 11x more efficient training ? They ca…

> The panic around deepseek is getting completely disconnected from reality.

Couldn’t agree more! Nobody here read the manual. The last paragraph of DeepSeek’s R1 paper:

> Software Engineering Tasks: Due to the long evaluation times, which impact the efficiency of the RL process, large-scale RL has not been applied extensively in software engineering tasks. As a result, DeepSeek-R1 has not demonstrated a huge improvement over DeepSeek-V3 on software engineering benchmarks. Future versions will address this by implementing rejection sampling on software engineering data or incorporating asynchronous evaluations during the RL process to improve efficiency.

Just based on my evaluations so far, R1 is not even an improvement on V3 in terms of real world coding problems because it gets stuck in stupid reasoning loops like whether “write C++ code to …” means it can use a C library or has to find a C++ wrapper which doesn’t exist.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#192

Earlier quoted context omitted.

Ollama has been deliberately misrepresenting R1 distill models as "R1" for marketing purposes. A lot of "AI" influencers on social media are unabashedly doing the same. Ollama's default "R1" model is a 4-bit RTN quantized 7B model, which is nowhere close to the real R1 (a 671B parameter fp8 MoE). https://www.reddit.com/r/LocalLLaMA/comments/1i8ifxd/ollama_...

Ollama is pretty clear about it, it's not like they are trying to deceive. You can also download the 671B model with Ollama, if you like.

no they are not, they intentionally remove every reference to this not being r1 from the cli and changed the names from the ones both Deepseek and Huggingface used.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#193
post #85
post #64

Earlier quoted context omitted.

> ran whatever version Ollama downloaded on a 3070ti (laptop version). It's reasonably fast. Probably was not r1, but one of the other models that got trained on r1, which apparently might still be quite good.

I'm not too hip to all the LLM terminology, so maybe someone can make sense of this and see if it's r1 or something based on r1: >>> /show info Model architecture qwen2 parameters 7.6B context length 131072 embedding length 3584 quantization Q4_K_M

it’s a distill, it’s going to be much much worse than r1

Re: Run DeepSeek R1 Dynamic 1.58-bit

#194
post #170

Hi small comment, please remember in china many things are sponsored by or subsidized by the government. "We[china] can do it for less.." , "it's cheaper in china.." only means the government gave us a pile of cash and help to get here . I 100% expect some downvotes from the ccp.

> Hi small comment, please remember in china many things are sponsored by or subsidized by the government. "We[china] can do it for less.." , "it's cheaper in china.." only means the government gave us a pile of cash and help to get here . And that's a really important strategic advantage China has versus America, which has such an insane fixation on pure(ish) free markets and free trade that it gives away its advant…

> And that's a really important strategic advantage China has versus America, which has such an insane fixation on pure(ish) free markets and free trade that it gives away its advantages in strategic industry after strategic industry.

> Some people falsely infer from the experience with the Soviet Union that freer markets always win geopolitical competition, but that's false.

The data we have is 500 years of free markets in the western world and the verdict is overwhelmingly: Yes, more freedom means more winning.

Just invite some incompetent bureaucrat over your house to dictate how you should cook and you'll quickly agree.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#195
post #161

Earlier quoted context omitted.

My experience gels with yours. Given the same code sample, DeepSeek has better, more creative suggestions about how to improve it, but it can't implement them without breaking the code. o1, generally, can implement DeepSeek's suggestions successfully. I think chaining them together might have quite interesting results.

Is there a tool that can automate chaining like that?

Aider has an architect mode where it asks one model to plan out the changes and another to actually write the code.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#196
post #170

Hi small comment, please remember in china many things are sponsored by or subsidized by the government. "We[china] can do it for less.." , "it's cheaper in china.." only means the government gave us a pile of cash and help to get here . I 100% expect some downvotes from the ccp.

> Hi small comment, please remember in china many things are sponsored by or subsidized by the government. "We[china] can do it for less.." , "it's cheaper in china.." only means the government gave us a pile of cash and help to get here . And that's a really important strategic advantage China has versus America, which has such an insane fixation on pure(ish) free markets and free trade that it gives away its advant…

It’s false except for every time that it has been true

Re: Run DeepSeek R1 Dynamic 1.58-bit

#197

An 80% size reduction is no joke, and the fact that the 1.58-bit version runs on dual H100s at 140 tokens/s is kind of mind-blowing. That said, I’m still skeptical about how practical this really is for most people. Like, yeah, you can run it on 24GB VRAM or even with just 20GB RAM, but "slow" is an understatement—those speeds would make even the most patient person throw their hands up. And then there’s the whole re…

People would only be 'throwing their hands up' because commercial LLMs have set unreasonable expectations for folks.

Anyone who has a/the need for or understands the value of a local LLM would be OK with this kind of output.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#198
post #34

As someone who is out of the loop, what’s the verdict on R1? Was anyone able to reproduce the results yet? Is the claim that it only took $5M to train generally accepted? It’s a very bold claim which is really shaking up the markets, so I can’t help but wonder if it was even verified at this point.

That's likely only the marginal cost of training this model, and doesn't include a lot of other costs, like the datacenters and GPUs themselves which they already had and also the staff. If they aren't lying because they have hardware they're not supposed to have, which is also a possibility.

these claims are getting more wrong every time i see them, weird game of telephone going around tech circles.

the cost absolutely includes the cost of GPUs and data centers, they quoted a standard price for renting h800 which has all of this built in. but yes, as very explicitly noted in the paper, it does not include cost of test iterations

Re: Run DeepSeek R1 Dynamic 1.58-bit

#199
post #168
post #74

Earlier quoted context omitted.

That was not the question.

It is if you're using market movements as evidence of anything factual. If markets aren't rational, you can't use them that way.

do you only take advice/learn from all-knowing people?

Re: Run DeepSeek R1 Dynamic 1.58-bit

#200
post #197

An 80% size reduction is no joke, and the fact that the 1.58-bit version runs on dual H100s at 140 tokens/s is kind of mind-blowing. That said, I’m still skeptical about how practical this really is for most people. Like, yeah, you can run it on 24GB VRAM or even with just 20GB RAM, but "slow" is an understatement—those speeds would make even the most patient person throw their hands up. And then there’s the whole re…

People would only be 'throwing their hands up' because commercial LLMs have set unreasonable expectations for folks. Anyone who has a/the need for or understands the value of a local LLM would be OK with this kind of output.

Everyone has the need for on device LLM, if the response rate was fast!
Post reply on HN