Live data from Hacker News

The gap between open weights LLMs and closed source LLMs

blog.doubleword.ai

141–150 of 264 posts

Re: The gap between open weights LLMs and closed source LLMs

#141
post #25
post #11

Earlier quoted context omitted.

We need a SETI@Home but for model training

I think model training is pretty hard to do efficiently on a vastly distributed network. If the model cant fit into the VRAM of the node your performance becomes so bad its useless, so a distributed model could only be properly trained if the size of the model doesnt exceed the majority of the nodes VRAM sizes. Maybe there is a different way of doing training but this would be the only way I can see. And it would sti…

With current paradigms, yes. I'm hoping to see more focus on architectures that are more amenable to distributed training in the near future.

Re: The gap between open weights LLMs and closed source LLMs

#142

IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.

Yeah, but the biggest plus for open models is that they can never be taken away. In other words, whatever capabilities they reach (even if there will never be another model), those stay forever. That can't be said for API-based models where a provider can sunset models whenever they feel like (i.e. gpt5-mini will soon be gone, and replaced by a more expensive 5.4-mini, same for goog, etc). And there will always be in…

Is this a valid point when we live in an evolving world. Language changes, facts change etc. Or can everything can just be grabbed from webpages and stored in the context window?

Re: The gap between open weights LLMs and closed source LLMs

#143
post #125

IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.

We should address the elephant in the room. The problem with the future of open weight models is not they are created as a result of philanthropy by some private org . All of the top contenders are created by the Chinese government . I don’t think we should describe these companies as simply releasing these highly capable open weight models out of the goodness of their hearts

None of those companies are created by the Chinese government. They're obviously subject to the Chinese government, whose whims may change at any given moment, but as we're seeing at the moment, so are the American companies.

And while I don't have a very positive view of the Chinese government, last I checked, they haven't been dropping bombs on innocent schoolchildren recently.

Re: The gap between open weights LLMs and closed source LLMs

#144
post #31

Earlier quoted context omitted.

Yeah, but the biggest plus for open models is that they can never be taken away. In other words, whatever capabilities they reach (even if there will never be another model), those stay forever. That can't be said for API-based models where a provider can sunset models whenever they feel like (i.e. gpt5-mini will soon be gone, and replaced by a more expensive 5.4-mini, same for goog, etc). And there will always be in…

True, but the capabilities and knowledge of that model are also frozen in time, so the value of that model declines over time. A model that writes code without knowledge of any language or library changes for half a decade is less useful. A 2021 era chatgpt would be quite quaint in 2026. Right now the Chinese labs might have incentives to release their models for free, and maybe Google is happy to release open weight…

If the world ends all I have to do is power my desktop and I'll have my locals - a decent iteration of Deepseek and a few smaller models, some focused, some just older versions - having several is key tho. They can be cross referenced to limit hallucinations and inaccurate information - this means I can confidently say that I have on my desktop - all of human history, knowledge, discoveries, maths, languages - at least in summary or truncated form (also another bonus of multiple models - will often have more comprehensive total output than one model provides) and all of those models have absolutely no restrictions other than the broadest limits allowed by current laws - so, practically no limits (I bet I could get them all it to explain splitting the atom with minor effort).

I realize that my amazing tool/system of local AI is out of date - I still very much like having it and it is not at all a bad thing to hav. Everyone in theory ought to have a local backup - for just in case.

The fact that people will have this in this one, albeit extreme, example - it would most definitely matter in the event of a societal collapse. Not everyone will have it - can they run those giant data centers off a few solar panels like a desktop PC?

For this one existential reason alone, I recommend everyone at least play around local enough to have a few models functional.

Re: The gap between open weights LLMs and closed source LLMs

#145
post #33

I believe the open model party will eventually end. Perhaps because companies realize it’s too much of a commercial advantage, countries don’t want to give other countries commercial or military help, or maybe even an outright ban after someone uses an open model to guide them through how to make a bomb.

Possible. Though I think open source innovation of hobbyists is currently hampered, because of open weights releases by chinese labs. I believe once the labs stop doing this, a globally coordinated open source ecosystem will fill the gap.

Re: The gap between open weights LLMs and closed source LLMs

#146
post #87

If the Chinese government is as involved in LLM development strategy as many people claim, wouldn't you expect them to immediately cease releasing open weight models and restrict access as soon as they start producing the frontier models? I am assuming this is what the USG thinks and is why they are trying to cut off the flow to foreign nationals ASAP. LLMs are an undeniably valuable tool, and governments like to con…

How do you know that Chinese don’t have powerful private models already? Maybe they just allow opening the ”bad” models…

As far as we don't know, they might also have operational stargates large enough to let their starships pass through. Actually, every country might have that.

But what is impactless to the wider world will always be as significant as something that never existed.

Re: The gap between open weights LLMs and closed source LLMs

#147
post #11

Earlier quoted context omitted.

We need a SETI@Home but for model training

Here's a project trying that - https://nousresearch.com/nous-psyche

Also https://pluralis.ai/ has distributed training (though they reach limits within seconds).

And I think https://allenai.org/ has something like this, too.

Re: The gap between open weights LLMs and closed source LLMs

#148

IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.

I think the bigger issue is the ever increasing capital requirements, which may cause even the closed weight companies to fall away from the frontier, e.g. Google & Meta are barely hanging on. For Google it feels a bit existential to remain at the frontier, but even then they're barely there. I hope that we find ways of continuing to improve these models besides continuing to exponentially increase capex spend until…

Isn't another issue that most successful open models are distilled from closed models, but closed models are putting more and better safeguards against distillation?

Re: The gap between open weights LLMs and closed source LLMs

#149
post #125

Earlier quoted context omitted.

We should address the elephant in the room. The problem with the future of open weight models is not they are created as a result of philanthropy by some private org . All of the top contenders are created by the Chinese government . I don’t think we should describe these companies as simply releasing these highly capable open weight models out of the goodness of their hearts

None of those companies are created by the Chinese government. They're obviously subject to the Chinese government, whose whims may change at any given moment, but as we're seeing at the moment, so are the American companies. And while I don't have a very positive view of the Chinese government, last I checked, they haven't been dropping bombs on innocent schoolchildren recently.

Bombs are a bit of a non sequitur here. The point is that Chinese companies are demonstrably hostile to American ones historically (and threatening in some specific structural ways to the American consumer). The presentation may be similar but to attribute American ethics to a Chinese decision is dubious.

Re: The gap between open weights LLMs and closed source LLMs

#150

IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.

> Until there's some sort of "community owned hardware"

The hardware is already available for renting at reasonable prices. We need community funding. I wish people pooled a fraction of the money they burn on local GPU rigs on funding training/testing/etc.

A big problem is like in open source: it's way too atomized. Just one competitive ground-up community LLM would require tens of millions $. But who gets to pick?

IMHO the only chance is highly specialized and smaller LLMs instead. And this is still millions to train.

And remember LLMs are competitive for only a handful months.

Post reply on HN