Live data from Hacker News

The gap between open weights LLMs and closed source LLMs

blog.doubleword.ai

41–50 of 264 posts

Re: The gap between open weights LLMs and closed source LLMs

#41

IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.

> The spigot can be turned off at any time. True. And it's possible that this has already happened at Alibaba Qwen - at least for the smaller models that people had a chance of running at home (122B and smaller).

We'll see. The qwen team has always released a few close to sota but proprietary models in between tgeir open releases. We did get 3.6 35B and 27B so its not all set in stone yet.

Its higley unlikely we get another open llama model though after the llama4 flop, even if their muse spark seems pretty good.

Re: The gap between open weights LLMs and closed source LLMs

#42
post #27

It would be interesting to know how much of a boost the closed models companies are giving the open models. If the closed models stop improving will the progress of open models slow?

Why are we assuming only American labs can innovate? DeepSeek already innovated a lot in efficiency, for example.

It's really unclear how much innovation DeepSeek has actually done, vs training on frontier model conversations.

Re: The gap between open weights LLMs and closed source LLMs

#44
post #22

Earlier quoted context omitted.

> they can never be taken away Your right to 3d print whatever you want is about to be taken away (in California). What software you can run on your computer can already be restricted. Absolutely everything can be taken away. The simplest way to remove open models is probably to declare them a tool that terrorists could use. Crazy? Yes, the world is totally crazy these days.

That only affects people in California. Whereas Fable being shut down affects people all over the world.

There's also, importantly, a distinction between what are told we can no longer use, and what can actually be taken away.

Open source and open hardware can be called illegal by a government, but, if we collectively invest our energy into open alternatives, they can't be taken away in the same sense. I can build a RepRap printer and I can use a local AI model. It's on all of us to make sure that the open alternatives are viable, maybe in the current global political reality now more than ever.

Making something illegal isn't a disincentive for everyone. When they start banning books, some of us start assembling printing presses.

Re: The gap between open weights LLMs and closed source LLMs

#45

Earlier quoted context omitted.

> they can never be taken away Your right to 3d print whatever you want is about to be taken away (in California). What software you can run on your computer can already be restricted. Absolutely everything can be taken away. The simplest way to remove open models is probably to declare them a tool that terrorists could use. Crazy? Yes, the world is totally crazy these days.

Just like declaring piracy illegal stopped piracy and removed pirated materials from everyone's computers. Everything cannot, in fact, be taken away. Don't propagandize yourself. Some things, like information, are free. Not even China can prevent all its citizens from accessing Western internet. USGov simply does not have the resources to find and audit every hard drive and USB stick in the country for illegal files.…

Remote attestation?

Re: The gap between open weights LLMs and closed source LLMs

#46
post #11

IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.

We need a SETI@Home but for model training

This has been a (noble) goal of lots of different projects in the community for a long time. Federated learning projects like Flower have been chipping away at it for a long time. There are many many hurdles to be cleared before anything in this area is super feasible as an alternative, but I applaud everyone who works on it.

Re: The gap between open weights LLMs and closed source LLMs

#47

The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure mo…

> Chinese labs must entirely retool from harvesting frontier model data to producing the data systems and efforts to produce novel data

Even if your characterization is accurate, they could do this tomorrow and are not so myopic that they wouldn’t have thought about it. I don’t see this as a barrier, and I see a lot of the same underestimation of Asia that’s been happening for 50 years. There’s not some innate American advantage to building LLMs, and personally I think whatever head start the US has is going to be squandered on delays from the export control “to dangerous for release” LARPing we’re seeing.

Re: The gap between open weights LLMs and closed source LLMs

#48

IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.

[deleted]

Re: The gap between open weights LLMs and closed source LLMs

#49

The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure mo…

Chinese frontier models don't need to catch up in every category. They just need to win in coding and that's exactly where they are going. The gap went from 12+ months to 1-2 months with the latest release of GLM 5.2 and coding is a task that you don't need heroic efforts to find rare and long-tail training data, you can just outsmart your competitor by optimizing algorithms and training recipes. This is something they can do at scale with the money and talent pool.

Re: The gap between open weights LLMs and closed source LLMs

#50
post #27

Earlier quoted context omitted.

Why are we assuming only American labs can innovate? DeepSeek already innovated a lot in efficiency, for example.

It's really unclear how much innovation DeepSeek has actually done, vs training on frontier model conversations.

Wym it's unclear? They publish their research...
Post reply on HN