Earlier quoted context omitted.
Or perhaps a 512GB Mac Studio. 671B Q4 of R1 runs on it.
I wouldn’t say runs. More of a gentle stroll.
DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
91–100 of 485 posts
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#92Earlier quoted context omitted.
I think the investment race here is an "all-pay auction"*. Lots of investors have looked at the ultimate prize — basically winning something larger than the entire present world economy forever — and think "yes". But even assuming that we're on the right path for that (which we may not be) and assuming that nothing intervenes to stop it (which it might), there may be only one winner, and that winner may not have even…
> investors have looked at the ultimate prize — basically winning something larger than the entire present world economy This is what people like Altman want investors to believe. It seems like any other snake oil scam because it doesn't match reality of what he delivers.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#93Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#94It's awesome that stuff like this is open source, but even if you have a basement rig with 4 NVIDIA GeForce RTX 5090 graphic cards ($15-20k machine), can it even run with any reasonable context window that isn't like a crawling 10/tps? Frontier models are far exceeding even the most hardcore consumer hobbyist requirements. This is even further
You can run at ~20 tokens/second on a 512GB Mac Studio M3 Ultra: https://youtu.be/ufXZI6aqOU8?si=YGowQ3cSzHDpgv4z&t=197 IIRC the 512GB mac studio is about $10k
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#95Disclaimer: I did not test this yet. I don't want to make big generalizations. But one thing I noticed with chinese models, especially Kimi, is that it does very well on benchmarks, but fails on vibe testing. It feels a little bit over-fitting to the benchmark and less to the use cases. I hope it's not the same here.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#96Disclaimer: I did not test this yet. I don't want to make big generalizations. But one thing I noticed with chinese models, especially Kimi, is that it does very well on benchmarks, but fails on vibe testing. It feels a little bit over-fitting to the benchmark and less to the use cases. I hope it's not the same here.
What is "Vibe testing"?
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#97Earlier quoted context omitted.
Home rigs like that are no longer cost effective. You're better off buying an rtx pro 6000 outright. This holds both for the sticker price, the supporting hardware price, the electricity cost to run it and cooling the room that you use it in.
I was just watching this video about a Chinese piece of industrial equipment, designed for replacing BGA chips such as flash or RAM with a good deal of precision: https://www.youtube.com/watch?v=zwHqO1mnMsA I wonder how well the aftermarket memory surgery business on consumer GPUs is doing.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#98Earlier quoted context omitted.
As much I agree with your sentiment, but I doubt the intention is singular.
I don't care if this kills Google and OpenAI. I hope it does, though I'm doubtful because distribution is important. You can't beat "ChatGPT" as a brand in laypeople's minds (unless perhaps you give them a massive "Temu: Shop Like A Billionaire" commercial campaign). Closed source AI is almost by design morphing into an industrial, infrastructure-heavy rocket science that commoners can't keep up with. The companies p…
Even when the technical people understood that, it would be too much of a political quagmire within their company when it became known to the higher ups. It just isn’t worth the political capital.
They would feel the same way about using xAI or maybe even Facebook models.
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#99Earlier quoted context omitted.
You can run at ~20 tokens/second on a 512GB Mac Studio M3 Ultra: https://youtu.be/ufXZI6aqOU8?si=YGowQ3cSzHDpgv4z&t=197 IIRC the 512GB mac studio is about $10k
and can be faster if you can get an MOE model of that
(commentary: things are really moving too fast for the layperson to keep up)
Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
#100Earlier quoted context omitted.
Two aspects to consider: 1. Chinese models typically focus on text. US and EU models also bear the cross of handling image, often voice and video. Supporting all those is additional training costs not spent on further reasoning, tying one hand in your back to be more generally useful. 2. The gap seems small, because so many benchmarks get saturated so fast. But towards the top, every 1% increase in benchmarks is sign…
Nothing you said helps with the issue of valuation. Yes, the US models may be better by a few percentage points, but how can they justify being so costly, both operationally as well as in investment costs? Over the long run, this is a business and you don't make money being the first, you have to be more profitable overall.