Live data from Hacker News

DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

huggingface.co

331–340 of 485 posts

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#331

Earlier quoted context omitted.

As much I agree with your sentiment, but I doubt the intention is singular.

It's like AMD open-sourcing FSR or Meta open-sourcing Llama. It's good for us, but it's nothing more than a situational and temporary alignment of self-interest with the public good. When the tables turn (they become the best instead of 4th best, or AMD develops the best upscaler, etc), the decision that aligns with self-interest will change, and people will start complaining that they've lost their moral compass.

>situational and temporary alignment of self-interest with the public good

That's how it supposed to work.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#333

How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…

> What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors

Quality was rarely the reason open source lagged in certain domains. Most of the time, open source solutions were technically superior. What actually hurt open source were structural forces, distribution advantages, and enterprise biases.

One could make an argument that open source solutions often lacked good UX historically, although that has changed drastically the past 20 years.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#334
post #325

Earlier quoted context omitted.

I don't know that it's actually prohibited. There is no Chinese telecommunications equipment allowed, no Huawei or Bytedance, but nothing prohibiting software merely being developed in China, not yet at least. Although I did just check what regions AWS bedrock support Deepseek and their govcloud regions do not, so that's a good reason not to use it. Still, on prem on a segmented network, following CMMC, probably perm…

> I don't know that it's actually prohibited. Chinese models generally aren't but DeepSeek specifically is at this point.

[deleted]

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#335

What version is actually running on chat.deepseek.com? It refuses to tell me when asked, only that it's been train with data up until July 2024, which would make it quite old. I turned off search and asked it for the winner of the US 2024 election, and it said it didn't know, so I guess that confirms it's not a recent model.

You can read that 3.2 is live on web and app here: https://api-docs.deepseek.com/news/news251201

The pdf describes how they did "continued pre-training" and then post training to make 3.2. I guess what's missing is the full pre-training that absorbs most date sensitive knowledge. That's probably also the reason that the versions are 3.x still.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#336

Why are there so few 32,64,128,256,512 GB models which could run on current consumer hardware? And why is the maximum RAM on Mac studio M4 128 GB??

As LLMs are productionised/commodified they're incorporating changes which are enthusiast-unfriendly. Small dense models are great for enthusiasts running inference locally, but for parallel batched inference MoE models are much more efficient.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#337
post #262

Earlier quoted context omitted.

Sparse Attention, it's the highlight of this model as per the paper

How did we come to the place that the most transparent and open models are now coming out of China—freely sharing their research and source code—while all the American ones are fully locked down

> How did we come to the place that the most transparent and open models are now coming out of China—freely sharing their research and source code—while all the American ones are fully locked down

Greed and "safety" hysteria.

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#338

Earlier quoted context omitted.

and can be faster if you can get an MOE model of that

All modern models are MoE already, no?

That's not the case. Some are dense and some are hybrid.

MOE is not the holy grail, as there are drawbacks eg. less consistency, expert under/over-use

Re: DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]

#340

How will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with acce…

> What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors Quality was rarely the reason open source lagged in certain domains. Most of the time, open source solutions were technically superior. What actually hurt open source were structural forces, distribution advantages, and enterprise biases. One could make an argument that open source solution…

For most professional software, the open source options are toys. Is there anything like an open source DAW, for example? It's not because music producers are biased against open source, it's because the economics of open source are shitty unless you can figure out how to get a company to fund development.
Post reply on HN