Live data from Hacker News

Run DeepSeek R1 Dynamic 1.58-bit

unsloth.ai

291–300 of 346 posts

Re: Run DeepSeek R1 Dynamic 1.58-bit

#291

Big fan of unsloth, they have huge potential, could definitely need some experienced GTM people though, IMO. The pricing page and messages sent there are really not good.

Oh thanks :) Yes agreed we do need better GTM - temporarily it's still me and my brother running Unsloth, so for now we're just prioritizing many more engineering releases :)

Re: Run DeepSeek R1 Dynamic 1.58-bit

#292

Earlier quoted context omitted.

> The panic around deepseek is getting completely disconnected from reality. This entire hype cycle has long been completely disconnected from reality. I've watched a lot of hype waves, and I've never seen one that oscillates so wildly. I think you're right that OpenAI isn't as hurt by DeepSeek as the mass panic would lead one to believe, but it's also true that DeepSeek exposes how blown out of proportion the initia…

In a sense it doesn't, in that if DeepSeek can do this , making OpenAI-type capabilities available for Llama-type infrastructure costs, then if you apply OpenAI scale infrastructure again to a much more efficient training/evaluation system, everything multiplies back up. I think that's where they'll have to head: using their infrastructure moat (such as it is) to apply these efficiency learnings to allow much more ca…

[deleted]

Re: Run DeepSeek R1 Dynamic 1.58-bit

#293

Earlier quoted context omitted.

At these prices, I would just get 2xDigits for $6k and have 256gb.

is it confirmed that you can get 256gb of vram for that amount? Because my understanding is that digits pricing will start at $3k for some basic config.

What they meant is buying two whole separate computers.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#294

Earlier quoted context omitted.

It absolutely doesn't. It sounds like further diluting the term "open-source" isn't great.

I assume when people say "open source model" they mean "open weights model". The "open source" term doesn't really make sense here, since machine learning models are not compilations of source code. (Though DeepSeek has published several papers with details on their training process. It's more than just open weights.)

ML models do have a "source" though

Re: Run DeepSeek R1 Dynamic 1.58-bit

#295

Earlier quoted context omitted.

Everyone has the need for on device LLM, if the response rate was fast!

I have MLCCHAT on my old Note 9 phone. It is actually still a great phone, but has 5GB RAM. Running an on device model is the first and only use case the RAM actually matters. And it has a headphone jack, OK? I just hate Bluetooth earbuds. And yeah, it isna problem, but I digress. When I run a 2.5B model, I get respectable output. Takes a minute or two to process the context, then output begins at somewhere on the or…

> First aid, how to make fires, materials and uses

This scares me more than it should...

Please do not trust an AI in actual life and death situations... Sure if it is literally your only option, but this implies you have a device on you that could make a phone call to an emergency number where a real human with real training and actually correct knowledge can assist you.

Even as an avid hiker the amount of times I've been out off cell service is miniscule and I absolutely refresh my knowledge on first aid regularly and any potential threats before a hike somewhere new.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#296

Earlier quoted context omitted.

I assume when people say "open source model" they mean "open weights model". The "open source" term doesn't really make sense here, since machine learning models are not compilations of source code. (Though DeepSeek has published several papers with details on their training process. It's more than just open weights.)

ML models do have a "source" though

If ML models have a source, brains have a source.

Brains don't have a source.

Therefore, ML models don't have a source.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#297

Earlier quoted context omitted.

is it confirmed that you can get 256gb of vram for that amount? Because my understanding is that digits pricing will start at $3k for some basic config.

What they meant is buying two whole separate computers.

I understand. It is still unclear if you can get 128GB vram for $3k.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#298
If I invested in a 100x machine because I needed 100 of x to run, and somebody shows how 10x can work, why have I not just become the holder of 10 10x machines, and therefore have already achieved capex to exploit this new market?

I cannot understand why "openai is dead" has legs: repurpose the hardware and data and it can be multiple instances of the more efficient model.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#299

An 80% size reduction is no joke, and the fact that the 1.58-bit version runs on dual H100s at 140 tokens/s is kind of mind-blowing. That said, I’m still skeptical about how practical this really is for most people. Like, yeah, you can run it on 24GB VRAM or even with just 20GB RAM, but "slow" is an understatement—those speeds would make even the most patient person throw their hands up. And then there’s the whole re…

> I’d rather build a rig with used 3090s and get way more bang for my buck

I'm curious, what would you use that rig for?

Re: Run DeepSeek R1 Dynamic 1.58-bit

#300

Earlier quoted context omitted.

You want to bet? The panic around deepseek is getting completely disconnected from reality. Don’t get me wrong what DS did is great, but anyone thinking this reshape the fundamental trend of scaling laws and make compute irrelevant is dead wrong. I’m sure OpenAI doesn’t really enjoy the PR right now, but guess what OpenAI/Google/Meta/Anthropic can do if you give them a recipe for 11x more efficient training ? They ca…

> You want to bet? Why would anyone bet? They can just short the OpenAI / MS stocks, and see in a few months if they were right or not.

How is that any different from a bet?
Post reply on HN