Live data from Hacker News

The gap between open weights LLMs and closed source LLMs

blog.doubleword.ai

51–60 of 264 posts

Re: The gap between open weights LLMs and closed source LLMs

#51
post #25
post #11

Earlier quoted context omitted.

We need a SETI@Home but for model training

I think model training is pretty hard to do efficiently on a vastly distributed network. If the model cant fit into the VRAM of the node your performance becomes so bad its useless, so a distributed model could only be properly trained if the size of the model doesnt exceed the majority of the nodes VRAM sizes. Maybe there is a different way of doing training but this would be the only way I can see. And it would sti…

My understanding is that in addition to your comment and the development of a method to separate the training data for distributed learning, the latency/bandwidth of systems connected on the internet is a challenge, too. Information has to be sent around before and after any hypothetical number crunching.

Re: The gap between open weights LLMs and closed source LLMs

#52

The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure mo…

Coding a case where it's possible to programmatically generate large amounts of data relatively cheaply. China could realistically surpass the US in coding while still being behind in many other areas.

Re: The gap between open weights LLMs and closed source LLMs

#53

IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.

Training these models is not a "hardware" problem.

Re: The gap between open weights LLMs and closed source LLMs

#54

The gap is huge and im tired of reading these articles constantly

Are you talking about hosted vs the ones you can easily run locally? Because there are open models that require hundreds of gb of vram which are apparently pretty close.

Re: The gap between open weights LLMs and closed source LLMs

#55

IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.

How is this a complaint? Once you have the model, you have the model. Download DeepSeek-R1 671B and you have it. You might not get improvements in the future, just like you may not ever get a future release of an open source project. Is that an indictment of open source?

But consider the alternative. OpenAI and Anthropic can shut off your account or API key at any time for any reason. How is this better? You have way more security when you're running your own model.

Re: The gap between open weights LLMs and closed source LLMs

#56

The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure mo…

How so? You'll soon have your choice of a very old OAI model or a new Chinese model, because the USG has no interest in letting you access the newest models without explicit permission.

Re: The gap between open weights LLMs and closed source LLMs

#57
post #56

The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure mo…

How so? You'll soon have your choice of a very old OAI model or a new Chinese model, because the USG has no interest in letting you access the newest models without explicit permission.

[deleted]

Re: The gap between open weights LLMs and closed source LLMs

#58
post #56

The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure mo…

How so? You'll soon have your choice of a very old OAI model or a new Chinese model, because the USG has no interest in letting you access the newest models without explicit permission.

Their point is that the Chinese models will also me limited to the very old OAI models, unless things flip. as they said.

The use of US models for Chinese model training is part of the motivation of all of this.

Re: The gap between open weights LLMs and closed source LLMs

#59
If the Chinese government is as involved in LLM development strategy as many people claim, wouldn't you expect them to immediately cease releasing open weight models and restrict access as soon as they start producing the frontier models? I am assuming this is what the USG thinks and is why they are trying to cut off the flow to foreign nationals ASAP.

LLMs are an undeniably valuable tool, and governments like to control those.

Re: The gap between open weights LLMs and closed source LLMs

#60
post #58
post #56

Earlier quoted context omitted.

How so? You'll soon have your choice of a very old OAI model or a new Chinese model, because the USG has no interest in letting you access the newest models without explicit permission.

Their point is that the Chinese models will also me limited to the very old OAI models, unless things flip . as they said. The use of US models for Chinese model training is part of the motivation of all of this.

Apologies - I was too quick in my response. I was speaking from a "how the users will perceive it" point of view. China's pretty good at the internet reputation thing.
Post reply on HN