The gap between open weights LLMs and closed source LLMs
71–80 of 264 posts
Re: The gap between open weights LLMs and closed source LLMs
#72[1] The story: https://nob.cs.ucdavis.edu/classes/ecs153-2019-04/readings/s...
[2] Wikipedia: https://en.wikipedia.org/wiki/Superiority_(short_story)
Re: The gap between open weights LLMs and closed source LLMs
#73If the Chinese government is as involved in LLM development strategy as many people claim, wouldn't you expect them to immediately cease releasing open weight models and restrict access as soon as they start producing the frontier models? I am assuming this is what the USG thinks and is why they are trying to cut off the flow to foreign nationals ASAP. LLMs are an undeniably valuable tool, and governments like to con…
Re: The gap between open weights LLMs and closed source LLMs
#74IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.
Plus I am certain it makes financial sense. I am guessing here but fully utilizing a subscriptions limits probably costs the operator more money than the subscription revenue, that is why anthropic is making such a big stink about the chinese data harvesting. By releasing the weights, you are relieving yourself from that burden because the competition does not need to hammer your subscription service they can just download your model and analyze it and run it all day.
Also for the largest models it makes no sense to run it yourself unless you are a major player. Renting the hardware is ludicrously more expensive than their subscription tens of thousands of dollars. And buying the hardware to run them is in the hundreds of thousands of dollars.
Re: The gap between open weights LLMs and closed source LLMs
#75Earlier quoted context omitted.
Why are we assuming only American labs can innovate? DeepSeek already innovated a lot in efficiency, for example.
It's really unclear how much innovation DeepSeek has actually done, vs training on frontier model conversations.
The Americans should wake up to reality because their fantasies that are repeated continuously in all Internet media, that supposedly the Chinese copy the US technology so they will not be able to surpass it, were true many years ago, but there are already many years since this theory has become false and now there are many domains where USA would have to copy the Chinese technology if they do not want to remain behind.
Among other "sanctions", USA has forbidden the export to China of high-performance computing devices, but this has backfired as China has just demonstrated a supercomputer that is faster than any US supercomputer and which uses custom CPUs designed in China, apparently by Huawei, the company that was the main target of the US efforts to sabotage the Chinese competitors.
The US "sanctions" have hurt China for a few years, but they have convinced them that they must allocate resources to become able to make themselves everything that they previously bought from USA. The result is that now China has become stronger and USA weaker.
USA should have never sold technology to China a quarter of century ago and then the power relationship between the 2 countries would have been very different. But even 5 years ago it was already too late for any US "sanctions" to have lasting effects. Nowadays any hopes that US "sanctions" will keep China in the dark ages are pathetic.
With the kind of policies that are promoted by the US government, the chances that USA will keep its leading position in AI are minimal.
Re: The gap between open weights LLMs and closed source LLMs
#76Earlier quoted context omitted.
Yeah, but the biggest plus for open models is that they can never be taken away. In other words, whatever capabilities they reach (even if there will never be another model), those stay forever. That can't be said for API-based models where a provider can sunset models whenever they feel like (i.e. gpt5-mini will soon be gone, and replaced by a more expensive 5.4-mini, same for goog, etc). And there will always be in…
> they can never be taken away Your right to 3d print whatever you want is about to be taken away (in California). What software you can run on your computer can already be restricted. Absolutely everything can be taken away. The simplest way to remove open models is probably to declare them a tool that terrorists could use. Crazy? Yes, the world is totally crazy these days.
Are laws that are inherently unenforceable even laws?
Re: The gap between open weights LLMs and closed source LLMs
#77Earlier quoted context omitted.
Yeah, but the biggest plus for open models is that they can never be taken away. In other words, whatever capabilities they reach (even if there will never be another model), those stay forever. That can't be said for API-based models where a provider can sunset models whenever they feel like (i.e. gpt5-mini will soon be gone, and replaced by a more expensive 5.4-mini, same for goog, etc). And there will always be in…
True, but the capabilities and knowledge of that model are also frozen in time, so the value of that model declines over time. A model that writes code without knowledge of any language or library changes for half a decade is less useful. A 2021 era chatgpt would be quite quaint in 2026. Right now the Chinese labs might have incentives to release their models for free, and maybe Google is happy to release open weight…
Re: The gap between open weights LLMs and closed source LLMs
#78The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure mo…
Distilling even with small amounts of data from a better model is still helpful, but not in the sense of transferring capabilities the raw internet-trained model doesn't have at all, but for identifying those capabilities that are compatible with the servile assistant persona and suppressing others that are undesirable (e.g. trolling). A primitive version of this were instruction-tuning datasets generated with ChatGPT, as used e.g. for Alpaca.
Without a clear target to emulate, competitors might have to rely more on human raters, but there are plenty of data labeling companies in China, so that's hardly a hurdle.
Re: The gap between open weights LLMs and closed source LLMs
#79Earlier quoted context omitted.
This seems backwards. Access to Fable can be removed. I don't see how an open weight model can ever be put back into the bag though.
The model itself, sure; the comment is about the production of more advanced models (to keep open weights near the frontier).
Re: The gap between open weights LLMs and closed source LLMs
#80Interesting to consider this inline with recent us export bans, could the US be squandering its lead by giving the open source, largely Chinese labs catch up (in terms of model quality available to masses), will US labs be able to maintain the lead without users being able to use their latest models?