Earlier quoted context omitted.
We need a SETI@Home but for model training
I think model training is pretty hard to do efficiently on a vastly distributed network. If the model cant fit into the VRAM of the node your performance becomes so bad its useless, so a distributed model could only be properly trained if the size of the model doesnt exceed the majority of the nodes VRAM sizes. Maybe there is a different way of doing training but this would be the only way I can see. And it would sti…
The gap between open weights LLMs and closed source LLMs
141–150 of 264 posts
Re: The gap between open weights LLMs and closed source LLMs
#142IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.
Yeah, but the biggest plus for open models is that they can never be taken away. In other words, whatever capabilities they reach (even if there will never be another model), those stay forever. That can't be said for API-based models where a provider can sunset models whenever they feel like (i.e. gpt5-mini will soon be gone, and replaced by a more expensive 5.4-mini, same for goog, etc). And there will always be in…
Re: The gap between open weights LLMs and closed source LLMs
#143IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.
We should address the elephant in the room. The problem with the future of open weight models is not they are created as a result of philanthropy by some private org . All of the top contenders are created by the Chinese government . I don’t think we should describe these companies as simply releasing these highly capable open weight models out of the goodness of their hearts
And while I don't have a very positive view of the Chinese government, last I checked, they haven't been dropping bombs on innocent schoolchildren recently.
Re: The gap between open weights LLMs and closed source LLMs
#144Earlier quoted context omitted.
Yeah, but the biggest plus for open models is that they can never be taken away. In other words, whatever capabilities they reach (even if there will never be another model), those stay forever. That can't be said for API-based models where a provider can sunset models whenever they feel like (i.e. gpt5-mini will soon be gone, and replaced by a more expensive 5.4-mini, same for goog, etc). And there will always be in…
True, but the capabilities and knowledge of that model are also frozen in time, so the value of that model declines over time. A model that writes code without knowledge of any language or library changes for half a decade is less useful. A 2021 era chatgpt would be quite quaint in 2026. Right now the Chinese labs might have incentives to release their models for free, and maybe Google is happy to release open weight…
I realize that my amazing tool/system of local AI is out of date - I still very much like having it and it is not at all a bad thing to hav. Everyone in theory ought to have a local backup - for just in case.
The fact that people will have this in this one, albeit extreme, example - it would most definitely matter in the event of a societal collapse. Not everyone will have it - can they run those giant data centers off a few solar panels like a desktop PC?
For this one existential reason alone, I recommend everyone at least play around local enough to have a few models functional.
Re: The gap between open weights LLMs and closed source LLMs
#145I believe the open model party will eventually end. Perhaps because companies realize it’s too much of a commercial advantage, countries don’t want to give other countries commercial or military help, or maybe even an outright ban after someone uses an open model to guide them through how to make a bomb.
Re: The gap between open weights LLMs and closed source LLMs
#146If the Chinese government is as involved in LLM development strategy as many people claim, wouldn't you expect them to immediately cease releasing open weight models and restrict access as soon as they start producing the frontier models? I am assuming this is what the USG thinks and is why they are trying to cut off the flow to foreign nationals ASAP. LLMs are an undeniably valuable tool, and governments like to con…
How do you know that Chinese don’t have powerful private models already? Maybe they just allow opening the ”bad” models…
But what is impactless to the wider world will always be as significant as something that never existed.
Re: The gap between open weights LLMs and closed source LLMs
#147Earlier quoted context omitted.
We need a SETI@Home but for model training
Here's a project trying that - https://nousresearch.com/nous-psyche
And I think https://allenai.org/ has something like this, too.
Re: The gap between open weights LLMs and closed source LLMs
#148IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.
I think the bigger issue is the ever increasing capital requirements, which may cause even the closed weight companies to fall away from the frontier, e.g. Google & Meta are barely hanging on. For Google it feels a bit existential to remain at the frontier, but even then they're barely there. I hope that we find ways of continuing to improve these models besides continuing to exponentially increase capex spend until…
Re: The gap between open weights LLMs and closed source LLMs
#149Earlier quoted context omitted.
We should address the elephant in the room. The problem with the future of open weight models is not they are created as a result of philanthropy by some private org . All of the top contenders are created by the Chinese government . I don’t think we should describe these companies as simply releasing these highly capable open weight models out of the goodness of their hearts
None of those companies are created by the Chinese government. They're obviously subject to the Chinese government, whose whims may change at any given moment, but as we're seeing at the moment, so are the American companies. And while I don't have a very positive view of the Chinese government, last I checked, they haven't been dropping bombs on innocent schoolchildren recently.
Re: The gap between open weights LLMs and closed source LLMs
#150IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.
The hardware is already available for renting at reasonable prices. We need community funding. I wish people pooled a fraction of the money they burn on local GPU rigs on funding training/testing/etc.
A big problem is like in open source: it's way too atomized. Just one competitive ground-up community LLM would require tens of millions $. But who gets to pick?
IMHO the only chance is highly specialized and smaller LLMs instead. And this is still millions to train.
And remember LLMs are competitive for only a handful months.