Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

91–100 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#91
post #46
post #38

Earlier quoted context omitted.

Or perhaps a 512GB Mac Studio. 671B Q4 of R1 runs on it.

I wouldn’t say runs. More of a gentle stroll.

With quantization, converting it to an MOE model... it can be a fast walk

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#92
post #70

Earlier quoted context omitted.

I think the investment race here is an "all-pay auction"*. Lots of investors have looked at the ultimate prize — basically winning something larger than the entire present world economy forever — and think "yes". But even assuming that we're on the right path for that (which we may not be) and assuming that nothing intervenes to stop it (which it might), there may be only one winner, and that winner may not have even…

> investors have looked at the ultimate prize — basically winning something larger than the entire present world economy This is what people like Altman want investors to believe. It seems like any other snake oil scam because it doesn't match reality of what he delivers.

Yeah, this is basically financial malpractice/fraud.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#94
post #55
post #9

It's awesome that stuff like this is open source, but even if you have a basement rig with 4 NVIDIA GeForce RTX 5090 graphic cards ($15-20k machine), can it even run with any reasonable context window that isn't like a crawling 10/tps? Frontier models are far exceeding even the most hardcore consumer hobbyist requirements. This is even further

You can run at ~20 tokens/second on a 512GB Mac Studio M3 Ultra: https://youtu.be/ufXZI6aqOU8?si=YGowQ3cSzHDpgv4z&t=197 IIRC the 512GB mac studio is about $10k

and can be faster if you can get an MOE model of that

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#95
post #51

Disclaimer: I did not test this yet. I don't want to make big generalizations. But one thing I noticed with chinese models, especially Kimi, is that it does very well on benchmarks, but fails on vibe testing. It feels a little bit over-fitting to the benchmark and less to the use cases. I hope it's not the same here.

This is why I stopped bothering checking out these models and, funnily enough, grok.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#96
post #51

Disclaimer: I did not test this yet. I don't want to make big generalizations. But one thing I noticed with chinese models, especially Kimi, is that it does very well on benchmarks, but fails on vibe testing. It feels a little bit over-fitting to the benchmark and less to the use cases. I hope it's not the same here.

What is "Vibe testing"?

He means capturing things that benchmarks don't. You can use Claude and GPT-5 back-to-back in a field that score nearly identically on. You will notice several differences. This is the "vibe".

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#97
post #30

Earlier quoted context omitted.

Home rigs like that are no longer cost effective. You're better off buying an rtx pro 6000 outright. This holds both for the sticker price, the supporting hardware price, the electricity cost to run it and cooling the room that you use it in.

I was just watching this video about a Chinese piece of industrial equipment, designed for replacing BGA chips such as flash or RAM with a good deal of precision: https://www.youtube.com/watch?v=zwHqO1mnMsA I wonder how well the aftermarket memory surgery business on consumer GPUs is doing.

LTT recently did a video on upgrading a 5090 to 96gb of ram

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#98
post #85

Earlier quoted context omitted.

As much I agree with your sentiment, but I doubt the intention is singular.

I don't care if this kills Google and OpenAI. I hope it does, though I'm doubtful because distribution is important. You can't beat "ChatGPT" as a brand in laypeople's minds (unless perhaps you give them a massive "Temu: Shop Like A Billionaire" commercial campaign). Closed source AI is almost by design morphing into an industrial, infrastructure-heavy rocket science that commoners can't keep up with. The companies p…

I can’t think of a single company I’ve worked with as a consultant that I could convince to use DeepSeek because of its ties with China even if I explained that it was hosted on AWS and none of the information would go to China.

Even when the technical people understood that, it would be too much of a political quagmire within their company when it became known to the higher ups. It just isn’t worth the political capital.

They would feel the same way about using xAI or maybe even Facebook models.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#99
post #55

Earlier quoted context omitted.

You can run at ~20 tokens/second on a 512GB Mac Studio M3 Ultra: https://youtu.be/ufXZI6aqOU8?si=YGowQ3cSzHDpgv4z&t=197 IIRC the 512GB mac studio is about $10k

and can be faster if you can get an MOE model of that

"Mixture-of-experts", AKA "running several small models and activating only a few at a time". Thanks for introducing me to that concept. Fascinating.

(commentary: things are really moving too fast for the layperson to keep up)

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#100

Earlier quoted context omitted.

Two aspects to consider: 1. Chinese models typically focus on text. US and EU models also bear the cross of handling image, often voice and video. Supporting all those is additional training costs not spent on further reasoning, tying one hand in your back to be more generally useful. 2. The gap seems small, because so many benchmarks get saturated so fast. But towards the top, every 1% increase in benchmarks is sign…

Nothing you said helps with the issue of valuation. Yes, the US models may be better by a few percentage points, but how can they justify being so costly, both operationally as well as in investment costs? Over the long run, this is a business and you don't make money being the first, you have to be more profitable overall.

[dead]
Post reply on HN