Live data from Hacker News

Run DeepSeek R1 Dynamic 1.58-bit

unsloth.ai

101–110 of 346 posts

Re: Run DeepSeek R1 Dynamic 1.58-bit

#101
The size reduction while keeping the model coherent is incredible. But I'm skeptical of how much effectiveness was retained. Flappy bird is well known and the kind of thing a non-reasoning model could het right. A better test would be something off the beaten path that R1 and o1 get right that other models don't.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#102
post #95
post #88

Earlier quoted context omitted.

Only on Hacker News would I have someone arguing with me that laptop theft is not a concern. You know what, you win. It's your $6,000 laptop, not mine.

So, zero times then. Ok!

For what it's worth, I never once insinuated that a laptop would get stolen on a train, only that I wouldn't want to bring such a laptop into the public in the first place. (Presumably, the laptop doesn't come into and exit existence upon entering and exiting the train, so this remains somewhat of a concern even if trains are involved.)

But yes, you're right. I've never personally seen a laptop get stolen. In fact, most people who have their laptop get stolen never see their laptop get stolen either.

I have, however, had coworkers who've had their laptops stolen. Multiple times.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#103
post #83

Earlier quoted context omitted.

I feel any AI can fix those problems when they can finally act. The problem AIs cannot run or debug code, or even book a hotel for me. When that is solved and an AI can interact with the code like a human does, it can fix its problems like a human does.

Exactly! Why can’t LLMs run their own code?

Rampancy.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#105
post #73

Earlier quoted context omitted.

Laptops get stolen on a train? An enclosed, single-direction space that only occasionally allows you to exit between infrequent, long-distance stops? A thing that contains ticket inspectors and a literal guard? How many laptops have you personally seen be stolen on a train?

You mean a tight, enclosed, single-direction space, crowded with people who are tired, and/or trying to relax, and/or thinking about the destination, and/or otherwise not particularly focused after hours of travel; a thing that contains ticket inspectors that show up every now and then to check tickets , and from which passengers embark and disembark at dozens point along the length of the thing, simultaneously, with…

Pickpocketing is a very different proposition. They relying on a lack of awareness, taking your wallet and being long gone before you’ve even noticed. If someone steals your laptop from in front of you without you even noticing I’d suggest that one is on you.

FWIW I’ve used my laptop on the train plenty, I’ve never had anything stolen nor felt in any danger of it.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#106
post #85
post #64

Earlier quoted context omitted.

> ran whatever version Ollama downloaded on a 3070ti (laptop version). It's reasonably fast. Probably was not r1, but one of the other models that got trained on r1, which apparently might still be quite good.

I'm not too hip to all the LLM terminology, so maybe someone can make sense of this and see if it's r1 or something based on r1: >>> /show info Model architecture qwen2 parameters 7.6B context length 131072 embedding length 3584 quantization Q4_K_M

"Qwen2.5 is the large language model series developed by Qwen team, Alibaba Cloud."

And I think they, the DeepSeek team, finetunes Qwen 7b on DeepSeek. That is how I understood it.

Which apparently makes it quite good for a 7b model. But, again: if I understood it correctly, is still just qween and without the reasoning of DeepSeek.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#107
post #40
post #34

As someone who is out of the loop, what’s the verdict on R1? Was anyone able to reproduce the results yet? Is the claim that it only took $5M to train generally accepted? It’s a very bold claim which is really shaking up the markets, so I can’t help but wonder if it was even verified at this point.

> Is the claim that it only took $5M to train generally accepted? Based on Nvidia being down 18% yesterday I would say the claim is generally accepted.

I don’t see them as related. The market moves when there is money to be made. It’s only tangentially related to any kind of general sentiment.

“I don’t believe this, but I know others will, so I’m selling”

Re: Run DeepSeek R1 Dynamic 1.58-bit

#108

Earlier quoted context omitted.

For personal usage, does it matter though? In most places residential electricity is cheap compared to everything else. In a DC context I feel it matters a lot more compared to the capex.

1x 3090 (350W power limit) already makes it feel like I'm running a fan heater under my desk, 5x would be nuts.

I think the last time any of my computers had a case was back when I realized the pair of 900gx2 cards I was running was turning my computer into an easy bake.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#109
post #73
post #71

Earlier quoted context omitted.

I am very happy for you, but laptops get stolen in public in most countries.

Laptops get stolen on a train? An enclosed, single-direction space that only occasionally allows you to exit between infrequent, long-distance stops? A thing that contains ticket inspectors and a literal guard? How many laptops have you personally seen be stolen on a train?

Yes, all the time. It's happened to two people I know, in France and in the US.

People get up to use the bathroom or the cafe car, the laptop is left behind for ten minutes, one of the train stops is while they're away from their seat, and someone sees an opportunity, snags it, and gets off at the stop.

This is an actual thing. And if it's worth a thousand bucks then it's very much worth getting off at an earlier stop then you'd planned, and continuing your journey on the next train.

Ticket inspectors or guards are irrelevant. There isn't one in your car 99% of the time.

I don't why you're trying to argue laptop theft on trains in first-world countries isn't a thing. It absolutely is.

Re: Run DeepSeek R1 Dynamic 1.58-bit

#110

It is going to be truly fucking revolutionary if open-source models are and continue to be able to challenge the state of the art. My big philosophical concern is that AI locks Capital into an absolutely supreme and insurmountable lead over Labour, and into the hands of oligarchs, and the possibility of a future where that's not case feels amazing. It pleases me greatly that this has Trump riled up too, because I thi…

I have no doubt open source will catch up (it already has, eh?) at the end of the day, it's just creative / new iterations on what is ultimately the transformer architecture... the amount of "secret" moat-like stuff that OpenAI was doing was bound to be figured out or exceeded eventually, like everything in tech...

Not to make fun of OpenAI and the great work they've done but it's kinda like if I went out in the 90s and said I'm going to found a company to have the best REST APIs. You can always found a successful tech company, but you can't found a successful tech company on a technological architecture or pattern alone.

Post reply on HN