Live data from Hacker News

The gap between open weights LLMs and closed source LLMs

blog.doubleword.ai

61–70 of 264 posts

Re: The gap between open weights LLMs and closed source LLMs

#61

Earlier quoted context omitted.

Just like declaring piracy illegal stopped piracy and removed pirated materials from everyone's computers. Everything cannot, in fact, be taken away. Don't propagandize yourself. Some things, like information, are free. Not even China can prevent all its citizens from accessing Western internet. USGov simply does not have the resources to find and audit every hard drive and USB stick in the country for illegal files.…

Remote attestation?

On PCs, the best you could really do is restrict access to certain websites on certain boxes with TPMs the users can't disable. Remote attestation can lock people out of your stuff, but not out of their own stuff. For that you need control of the device. Of course, most mobile phones aren't easy for the user to have control of, but most PCs still are, so long as you scrub the rootkits (e.g. windows) off 'em

Re: The gap between open weights LLMs and closed source LLMs

#62

The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure mo…

“China can only copy the US” is a very short sighted and uninformed opinion. there is more coming out of china than just new ways to distill models

Re: The gap between open weights LLMs and closed source LLMs

#63
post #47

The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure mo…

> Chinese labs must entirely retool from harvesting frontier model data to producing the data systems and efforts to produce novel data Even if your characterization is accurate, they could do this tomorrow and are not so myopic that they wouldn’t have thought about it. I don’t see this as a barrier, and I see a lot of the same underestimation of Asia that’s been happening for 50 years. There’s not some innate Americ…

Why would those have any impact on R&D speed? Most are funded and close to cash flow positive

Re: The gap between open weights LLMs and closed source LLMs

#64

IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.

Training these models is not a "hardware" problem.

[deleted]

Re: The gap between open weights LLMs and closed source LLMs

#65

IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.

Training these models is not a "hardware" problem.

I think that simplifies it a bit. You can't train without hardware, which is why the Chinese companies are illegally importing Nvidia cards [1].

[1] https://www.theinformation.com/articles/deepseek-using-banne...

Re: The gap between open weights LLMs and closed source LLMs

#66

The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure mo…

Unless you are working at one of these companies you don't know what they are doing.

You don't know what's happening in z.ai nor alibaba. And you don't know what's happening in anthropic and open ai.

I don't know what they are all doing, but I find it extremely unlikely that they are not all collecting data from one another. I am confident anthropic has a team going over GML 5.2 weights even if it's just to see where the competition is.

Just because some labs are getting data from Anthropic does not mean they are not also doing their own research.

They were focused on optimization because they could not get the best hardware.The only reason their top labs are behind may be because they did not have h200s and MI350s. And now they do.

Plus you are discounting other risks, Anthropic is currently sitting on "the best" models in the world because they got in a pissing match with the US administration.

btw: This could be the case in china as well, their administration has been surprisingly open on AI exports and open weight models, that we know of. There is a very small but not trivial chance they are hogging a better version of glm 5.2 for example, but no one is allowed to talk about it. Now I am not saying that is the case, I am saying the two cases (chinese labs are 6 months behind, they are forced to suppress their best models) are indistinguishable.

Re: The gap between open weights LLMs and closed source LLMs

#67

The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure mo…

Yeah, this is, to be perfectly blunt, cope, for several reasons:

1. It's unclear if there is a law of diminishing returns with ever-larger models. They're more expensive to run and for many applications, you'll probably find smaller models are sufficient;

2. There's an inbuilt market for local LLMs. This is an effective limit on how large models can get. Case law hasn't been established yet on, for example, if a law firm using ChatGPT breaks privilege. Specifically, chat logs may be discoverable. Medical applications have this issue too and I think you'll find that financial firms are going to be leery about this as well;

3. Better, larger models will bleed into smaller, open source models. The chat logs themselves are training data. There's a whole market in China for Claude tokens around this;

4. China has a national security interest in not being beholden to US tech giants when it comes to AI. China has a history of being able to commit to large-scale long-term projects and Anthropic just won't be able to compete with a national project by one of the world's superpowers, if it comes down to it;

5. Winning doesn't necessarily mean being the best. Often it's just being good enough;

6. As an example of a national project, China is busy replicating EUV because of the US ban on ASML and NVidia exporting their best stuff. I don't think many in the West are prepared for how rapid this will be. I'm reminded of the policy debate in 1945 when many in American policy and militarey circles thought the USSR would never catch up with atomic bomb or, if they did, it would take 20+ years. It took 4 years. For the hydrogen bomb, it took 1. The US hardware advantage is a lot more tenuous than many realize.

Re: The gap between open weights LLMs and closed source LLMs

#68
post #47

The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure mo…

> Chinese labs must entirely retool from harvesting frontier model data to producing the data systems and efforts to produce novel data Even if your characterization is accurate, they could do this tomorrow and are not so myopic that they wouldn’t have thought about it. I don’t see this as a barrier, and I see a lot of the same underestimation of Asia that’s been happening for 50 years. There’s not some innate Americ…

Exactly. If they wanted to they could produce the same amount of data. Companies like Scale, Mercor, Surge exists for a reason, a reason that doesn't need to exist in China if they mandate Chinese enterprises to provide all their real world data (or have them work inside RL environments) to the model companies for post training. There is no real advantage that US companies have except a head start, and as Jensen said, a ton of the research advantage is skewed since a lot of the best researchers in the US are Chinese nationals. I do think the model is just one piece of the pie (not to echo Jensen too much), and hopefully we will always be able to serve these bigger frontier models in a much more efficient way as well as building out the application layer faster which actually makes them useful and/or more dangerous/powerful.

Re: The gap between open weights LLMs and closed source LLMs

#69
post #34

Frankly it does not matter if there is gap because for most practical use-cases the end user can barely perceive the difference in intelligence. On paper frontier models will be ahead of the curve but I don't think hardly anyone will be able to tell if a piece of work, say a landing page, is created with Fable or GLM and that is the point. The perceptible intelligence will reach a point beyond which it is no longer c…

> except for some narrow use-case.

I think it's entirely the opposite. For narrow use cases, like web pages and crud/GUI, the open source models don't show much of a difference.

Re: The gap between open weights LLMs and closed source LLMs

#70
post #49

The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure mo…

Chinese frontier models don't need to catch up in every category. They just need to win in coding and that's exactly where they are going. The gap went from 12+ months to 1-2 months with the latest release of GLM 5.2 and coding is a task that you don't need heroic efforts to find rare and long-tail training data, you can just outsmart your competitor by optimizing algorithms and training recipes. This is something th…

> They just need to win in coding and that's exactly where they are going.

They don't even need to 'win' in the sense of maxing the benchmark. They can be 20% worse/50% cheaper and many of us (and our managers who approve our token budgets) will be in.

Deepseek is 30x cheaper for input/75x cheaper for output than sonnet on openrouter, and it's not a whole lot worse for many things.

Post reply on HN