Earlier quoted context omitted.
>Like, I get that shared memory architectures like a 192GB Mac Ultra are a big deal, but who’s dropping $6,000+ on that setup? AMD strix halo APU will have quad channel memory and will launch soon so expect these kinds of setups available for much less. Apple is charging an arm and a leg for memory upgrades, hopefully we get competition soon. From what I saw at CES OEMs are paying attention to this use case as well -…
Unfortunately, Apple’s RAM and Storage upgrade prices are very in line with other class comparable OEMs. I’m sure there’ll be some amount of undercutting but I don’t think it’ll be a huge difference on the RAM side itself.
Run DeepSeek R1 Dynamic 1.58-bit
181–190 of 346 posts
Re: Run DeepSeek R1 Dynamic 1.58-bit
#182Earlier quoted context omitted.
I am very happy for you, but laptops get stolen in public in most countries.
Laptops get stolen on a train? An enclosed, single-direction space that only occasionally allows you to exit between infrequent, long-distance stops? A thing that contains ticket inspectors and a literal guard? How many laptops have you personally seen be stolen on a train?
Re: Run DeepSeek R1 Dynamic 1.58-bit
#183Wow, an 80% reduction in size for DeepSeek-R1 is just amazing! It's fantastic to see such large models becoming more accessible to those of us who don't have access to top-tier hardware. This kind of optimization opens up so many possibilities for experimenting at home. I'm impressed by the 140 tokens per second speed with the 1.58-bit quantization running on dual H100s. That kind of performance makes the model pract…
Btw completely off topic, but your comment triggered the internal classification in my brain, and it looks like AI-generated. Not accusing you anything. Could be that you happen to write in a way similar to LLMs. Could be that we are influenced by LLM writing styles and are writing more and more like LLMs. Could be that the difference between LLM generated content and human-generated content is getting smaller and ha…
It’s the exclamation point in the first paragraph, the concise and consistent sentence structure, and the lack of colloquial tone.
OP, no worries if you’re real. I often read my own messages or writing and worry that people will think I’m an LLM too.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#184Earlier quoted context omitted.
Yes, all the time. It's happened to two people I know, in France and in the US. People get up to use the bathroom or the cafe car, the laptop is left behind for ten minutes, one of the train stops is while they're away from their seat, and someone sees an opportunity, snags it, and gets off at the stop. This is an actual thing. And if it's worth a thousand bucks then it's very much worth getting off at an earlier sto…
Different regions of the world would see different degrees of responsibilities regarding theft. I would consider absurd to leave unattended in a public space something valuable, considering the effort required to avoid that (that is: taking it with you). So, yes, theft on trains for people that think they are 100% safe are a thing, but applying the same idea (to assume something is 100% safe and not be cautious) I wo…
I don't think I've ever seen a human being do that before on a train. Not to go to the toilet, nor to grab a coffee in another car.
You can't be paranoid about everything. My friend in France had put his laptop back into his bag where it wasn't visible and assumed that was good enough, but someone must have seen him do it and just took the whole bag.
You are applying a totally unreasonable standard, to suppose that the thefts were due to unreasonable carelessness. What, do you think someone should take their large luggage into the bathroom too, every time they need to pee?
Talk about victim-blaming.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#185Re: Run DeepSeek R1 Dynamic 1.58-bit
#186Earlier quoted context omitted.
Honestly, if you have a residence of some kind and an Internet connection, you don't need to bring your beefy computer with you everywhere. It is cool to be able to have ridiculously powerful mobile computers, but I don't think I would ever be willing to take a $6,000 laptop anywhere it has a decent chance of being stolen.
Do you live in a third world country? If so I might agree, but otherwise trains are perfectly safe.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#187Earlier quoted context omitted.
> Can I use that on the train though? I can with a 128GB MacBook, without it sounding like a helicopter taking off as well. What kind of timescale do you expect to be able to train a useful LLM with that?
Well it’s about an hour to commute on the train so I guess that long :3
Re: Run DeepSeek R1 Dynamic 1.58-bit
#188As someone who is out of the loop, what’s the verdict on R1? Was anyone able to reproduce the results yet? Is the claim that it only took $5M to train generally accepted? It’s a very bold claim which is really shaking up the markets, so I can’t help but wonder if it was even verified at this point.
If they aren't lying because they have hardware they're not supposed to have, which is also a possibility.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#189Earlier quoted context omitted.
In my experience with deepseek and o1, openai's big talk about (and investment into) hallucination avoidance might save their hides here. Deepseek may be smarter, and understand complex problems better, but it also seems to make mistakes more often. (It's as if it's comprehension is better, but it's worse at memorization/recall.) Need an LLM to one-shot some complex network scripting? as of last night, o1 is still wh…
My experience gels with yours. Given the same code sample, DeepSeek has better, more creative suggestions about how to improve it, but it can't implement them without breaking the code. o1, generally, can implement DeepSeek's suggestions successfully. I think chaining them together might have quite interesting results.
Re: Run DeepSeek R1 Dynamic 1.58-bit
#190An 80% size reduction is no joke, and the fact that the 1.58-bit version runs on dual H100s at 140 tokens/s is kind of mind-blowing. That said, I’m still skeptical about how practical this really is for most people. Like, yeah, you can run it on 24GB VRAM or even with just 20GB RAM, but "slow" is an understatement—those speeds would make even the most patient person throw their hands up. And then there’s the whole re…
Not everyone needs the largest model. There are variations or R1 with fewer parameters that can easily run on consumer hardware. With 80% size reduction you could run 70B on 8-bit on an RTX 3090. Other than that, if you really need the big one you can get six 3090s and you're good to go. It's not cheap, but you're running a ChatGPT equivalent model from your basement. A year ago this was a wetdream for most enthusias…