Live data from Hacker News

Run DeepSeek R1 Dynamic 1.58-bit

unsloth.ai

211–220 of 346 posts

Re: Run DeepSeek R1 Dynamic 1.58-bit

#211

> DeepSeek-R1 has been making waves recently by rivaling OpenAI's O1 reasoning model while being fully open-source. Do we finally have a model with access to the training architecture and training data set, or are we still calling non-reproducible binary blobs without source form open-source?

It sounds like if they owe you the training architecture and training data set.

It absolutely doesn't. It sounds like further diluting the term "open-source" isn't great.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#212

Earlier quoted context omitted.

Pickpocketing is a very different proposition. They relying on a lack of awareness, taking your wallet and being long gone before you’ve even noticed. If someone steals your laptop from in front of you without you even noticing I’d suggest that one is on you. FWIW I’ve used my laptop on the train plenty, I’ve never had anything stolen nor felt in any danger of it.

But would you consider leaving it unattended on your seat and going for lunch to the restaurant car, or for an extended toilet break?

You might have seen some laptops have screens that fold down, I know MacBooks do. This "clam shell" effect protects the keyboard, trackpad, and even the screen from bumps and jostles. Many laptops when so closed can even fit in a backpack.

So a little trick I figured out is to close my laptop lid and then slide it into a pocket of my backpack. I can then carry it with me when I get up and move around.

So then I can take it with me to eat lunch or an extended toilet break. Maybe some day all laptops will have that feature.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#213

Earlier quoted context omitted.

You want to bet? The panic around deepseek is getting completely disconnected from reality. Don’t get me wrong what DS did is great, but anyone thinking this reshape the fundamental trend of scaling laws and make compute irrelevant is dead wrong. I’m sure OpenAI doesn’t really enjoy the PR right now, but guess what OpenAI/Google/Meta/Anthropic can do if you give them a recipe for 11x more efficient training ? They ca…

Computing is not king, DeepSeek just demonstrated otherwise. And yes, OpenAI will have to reinvent itself to copy DS, but this means they'll have to throw away a lot of their investment in existing tech. They might recover but it is not a minor hiccup as you suggest.

I just don't see how this is true. OpenAI has a massive cash & hardware pile -- they'll adapt and learn from what DeepSeek has done and be in a position to build and train 10x-50x-100x (or however) faster and better. They are getting a wake-up call for sure but I don't think much is going to be thrown away.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#216
post #197

An 80% size reduction is no joke, and the fact that the 1.58-bit version runs on dual H100s at 140 tokens/s is kind of mind-blowing. That said, I’m still skeptical about how practical this really is for most people. Like, yeah, you can run it on 24GB VRAM or even with just 20GB RAM, but "slow" is an understatement—those speeds would make even the most patient person throw their hands up. And then there’s the whole re…

People would only be 'throwing their hands up' because commercial LLMs have set unreasonable expectations for folks. Anyone who has a/the need for or understands the value of a local LLM would be OK with this kind of output.

I use commercial LLMs every day. The best of them can still be infuriating at times to the point of being unproductive. So I'm not sure I agree here.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#217
post #209

Earlier quoted context omitted.

Btw completely off topic, but your comment triggered the internal classification in my brain, and it looks like AI-generated. Not accusing you anything. Could be that you happen to write in a way similar to LLMs. Could be that we are influenced by LLM writing styles and are writing more and more like LLMs. Could be that the difference between LLM generated content and human-generated content is getting smaller and ha…

haha you got me. I'm real person using LLM to proofread the stuff I write. English is not my native language and I'm trying to improve my written vocabulary a little bit. Sorry if it reads a little bit too off.

Haha no worries. This is a perfectly valid use case of LLM. I'm happy that the comment sounds very professional and to the point.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#219

Earlier quoted context omitted.

> They do this so you'll have to buy several cards for your AI workstation. AFAIK you can't do that with newer consumer cards, which is why this became an annoyance. Even a RTX 4070 Ti with its 12 GB would be fine, if you could easily stack a bunch of them like you used to be able with older cards.

It's "easy" if you have a place to build an open frame rig with riser cables and whatnot. I can't do that, so I'm going the single slot waterblock route, which unfortunately rules out 3090s due to the memory on the back side of the PCB. It's very frustrating.

There was a rumor that 5090 or 5090D for China may or may not come with multi-GPU software locked. I think GP's referring to that. It's not clear if it is the case with retail cards.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#220
post #168

Earlier quoted context omitted.

It is if you're using market movements as evidence of anything factual. If markets aren't rational, you can't use them that way.

do you only take advice/learn from all-knowing people?

Do you know any?

But here's my advice: drop the fallacious arguments and try something more honest.

Post reply on HN