Live data from Hacker News

Run DeepSeek R1 Dynamic 1.58-bit

unsloth.ai

341–346 of 346 posts

Re: Run DeepSeek R1 Dynamic 1.58-bit

#341
post #91

Earlier quoted context omitted.

I canceled my OpenAI subscription last night, as did many many others. There were some threads in reddit with everyone chiming in they all just canceled too. imo OpenAI is done, and will go through massive cuts and probably acquired by the end of the year for a very tiny fraction of its current value.

You want to bet? The panic around deepseek is getting completely disconnected from reality. Don’t get me wrong what DS did is great, but anyone thinking this reshape the fundamental trend of scaling laws and make compute irrelevant is dead wrong. I’m sure OpenAI doesn’t really enjoy the PR right now, but guess what OpenAI/Google/Meta/Anthropic can do if you give them a recipe for 11x more efficient training ? They ca…

...until they distill that "11x efficient training" again ...

Re: Run DeepSeek R1 Dynamic 1.58-bit

#342
post #324

Earlier quoted context omitted.

> The data we have is 500 years of free markets in the western world and the verdict is overwhelmingly: Yes, more freedom means more winning. No, more freedom means more winning to a point . Past that point it does not, and I'd argue that's where the US is. > Just invite some incompetent bureaucrat over your house to dictate how you should cook and you'll quickly agree. That's supposed to be convincing, somehow? Just…

"I'm from the government and I'm here to help" are words you like to hear? At least anybody else I can tell to leave me alone.

> "I'm from the government and I'm here to help" are words you like to hear?

They're definitely not "the nine most terrifying words in the English language." Government is a necessity and performs important functions: we'd be worse off without it. A libertarian utopia would actually be a dystopia, at least for the vast majority.

Some day, historians will ask: Why did China eclipse the United States? And the tl;dr answer will likely be: libertarians. Myopic enthusiasm for free markets has really degraded the US's ability to make strategic decisions to maintain its advantages, and it seems on track to walk down the value chain while a few people get really rich leading it that way.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#343

Earlier quoted context omitted.

How'd it go, and which client are you using? :)

Pretty rough. Using LM Studio, trying to load the model throws an error of "insufficient system resources." I disabled this error, set the context length to 1024 and was able to get 0.24 tokens per second. Comparatively, the 32B distill model gets about 20 tokens per second. And it became incredibly flaky, using up all available ram, and crashing the whole system a few times. While the M4 Max 128GB handles the 32B we…

I had the same (compiled llama.cpp myself). Changed it to all CPU I think (num layers on GPU to 0) and it went up to 1.8 tokens per second. I think it can go up much more

Re: Run DeepSeek R1 Dynamic 1.58-bit

#344
post #342

Earlier quoted context omitted.

"I'm from the government and I'm here to help" are words you like to hear? At least anybody else I can tell to leave me alone.

> "I'm from the government and I'm here to help" are words you like to hear? They're definitely not "the nine most terrifying words in the English language." Government is a necessity and performs important functions: we'd be worse off without it. A libertarian utopia would actually be a dystopia, at least for the vast majority. Some day, historians will ask: Why did China eclipse the United States? And the tl;dr ans…

> Some day, historians will ask: Why did China eclipse the United States?

Why do you need to make leaps to the future to find evidence for your claims? Anyone can simply look at the past 500 years and come to the opposite conclusion.

There's also not much evidence China will "eclipse" the United States (whatever that means). I hitchhiked mainland China in 2019 after studying the language in university precisely because I thought the country might "eclipse" mine.

I came back with the exact opposite conclusion.

If the definition of "eclipse" is more global cultural influence, I would challenge you to compare the number of American movies you've watched in the past year vs Chinese movies. Movies are just 1 dimension of this dynamic.

The country has simply too much history and insularity to propagate its influence throughout the world. The language is another key example- very few learn Chinese as a second language. Even the Chinese youth themselves often use the latin alphabet to write their own language on a keyboard.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#345

Hi small comment, please remember in china many things are sponsored by or subsidized by the government. "We[china] can do it for less.." , "it's cheaper in china.." only means the government gave us a pile of cash and help to get here . I 100% expect some downvotes from the ccp.

https://apnews.com/article/deepseek-china-generative-ai-inte... TA

Re: Run DeepSeek R1 Dynamic 1.58-bit

#346
post #85
post #64

Earlier quoted context omitted.

> ran whatever version Ollama downloaded on a 3070ti (laptop version). It's reasonably fast. Probably was not r1, but one of the other models that got trained on r1, which apparently might still be quite good.

I'm not too hip to all the LLM terminology, so maybe someone can make sense of this and see if it's r1 or something based on r1: >>> /show info Model architecture qwen2 parameters 7.6B context length 131072 embedding length 3584 quantization Q4_K_M

Hi Kye, I tried a version of this model to assess its capabilities.

I would recommend you to try to run the llama-based distill (same size, same quantization) that you can find here: https://huggingface.co/bartowski/DeepSeek-R1-Distill-Llama-8...

It should take the same amount of memory as the one you currently have.

In my experience the Llama version performs much better at adhering to the prompt, understanding data in multiple languages, and going in-depth in its responses.

Post reply on HN