Live data from Hacker News

The gap between open weights LLMs and closed source LLMs

blog.doubleword.ai

221–230 of 264 posts

Re: The gap between open weights LLMs and closed source LLMs

#221
post #125

Earlier quoted context omitted.

We should address the elephant in the room. The problem with the future of open weight models is not they are created as a result of philanthropy by some private org . All of the top contenders are created by the Chinese government . I don’t think we should describe these companies as simply releasing these highly capable open weight models out of the goodness of their hearts

None of those companies are created by the Chinese government. They're obviously subject to the Chinese government, whose whims may change at any given moment, but as we're seeing at the moment, so are the American companies. And while I don't have a very positive view of the Chinese government, last I checked, they haven't been dropping bombs on innocent schoolchildren recently.

Go to chat.z.ai right now and ask it about what happened in Tiananmen Square. Do you think it's good for the world if software is written by the model that answers that question that way?

Re: The gap between open weights LLMs and closed source LLMs

#222
post #201

Earlier quoted context omitted.

Are you sure you haven't gotten your catastrophes crossed? Ozone depletion was a different crisis and people did enact change, the ozone hole has been closing fairly steadily. Wikipedia [0] thinks the prospects for the ozone layer are pretty good. [0] https://en.wikipedia.org/wiki/Ozone_depletion#Prospects_of_o...

If only climate change was as easy a solve

Ozone action occurred before the rise of social media.

We were fortunate. 20 years later and we’d have been screwed

Re: The gap between open weights LLMs and closed source LLMs

#223

IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.

I am the original author of the post - thanks for reading it! I think the future of open weights models will be similar to fabless chip design companies. There will be companies that can train models and they will licence those models to inference companies that manage the APIs. The inference companies need much less capital and the training companies dont need to divert resources from training to inference. Some of…

On mobile, or at least on mine (pixel 10) using chrome, the graphs are unreadable and unusable, which is a shame as I'm quite interested in them and I don't have access to a pc at the moment. Would you be able to change that?

They steal the scroll/drag touch and turn into a nightmare if zooming / unzooming, and are squashed and unreadable when they first render.

Re: The gap between open weights LLMs and closed source LLMs

#224

IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.

[flagged]

Re: The gap between open weights LLMs and closed source LLMs

#225

> Now is probably a good time to liquidate your pension, fly to a remote island somewhere, and live out the remaining 6 months or so of civilization in peace. > So maybe the open source apocalypse won’t happen yet. Sorry I wasn't at the last doomer meeting, when did we decide good open source models are a harbinger for the apocalypse?

I assumed it meant that when open weights reach the capability of frontier models, and tounge in cheek referencing the terrible consequences of us all getting our hands on mythos+ capability models without restrictions.

Re: The gap between open weights LLMs and closed source LLMs

#226

Achilles and the tortoise [0] is usually a fallacy. If the tortoise has a head start, then Achilles will never catch it because in the time it takes Achilles to reach the tortoise's location the tortoise has moved some degree further, ad infinitum. Obviously not real because Achilles will pass the tortoise -- I think a fallacy because the framing creates a fake asymptote (they will both pass the point where they're a…

comparing a thought experiment about relative movement through an alleged continuum over the sum of infinitesimal quasi-instants to the release cadence and maturation of open weights to proprietary LLMs is super bizarro guy

Re: The gap between open weights LLMs and closed source LLMs

#227

Earlier quoted context omitted.

Yeah, but the biggest plus for open models is that they can never be taken away. In other words, whatever capabilities they reach (even if there will never be another model), those stay forever. That can't be said for API-based models where a provider can sunset models whenever they feel like (i.e. gpt5-mini will soon be gone, and replaced by a more expensive 5.4-mini, same for goog, etc). And there will always be in…

> they can never be taken away Your right to 3d print whatever you want is about to be taken away (in California). What software you can run on your computer can already be restricted. Absolutely everything can be taken away. The simplest way to remove open models is probably to declare them a tool that terrorists could use. Crazy? Yes, the world is totally crazy these days.

> What software you can run on your computer can already be restricted.

Yes, over my dead body.

Re: The gap between open weights LLMs and closed source LLMs

#228

Earlier quoted context omitted.

on the Will It Mythos benchmark, small models are punching way above their weight(s) gemma4-26B (#7) qwen-3.6-27B (#9) https://news.ycombinator.com/item?id=48640196

I've tried running qwen 3.6 locally and it felt like LLMs a year ago where you can get them to do some stuff but the tasks have to be very small and you have to course correct them a lot to the point it's hard to say it's any faster than doing it all yourself. Certainly the gap is closing but I feel it still makes more sense to pay pennies to run the full sized open models hosted on much better hardware.

I had qwen36moe revamp my PhD thesis with a rewrite using JAX. Gave it access to my old code, helpednitnwhen it got stuck or didn't quite understand a few times.

Overall I was very impressed with its open box reimplementation. I remain of the mind they are widely underrated.

Re: The gap between open weights LLMs and closed source LLMs

#229

Earlier quoted context omitted.

Why would China care about deflating the US AI bubble? Why do we think there is a bubble for sure in 2025/2026? Why doesn't China also worry about their own AI bubble inside the country?

>Why would China care about deflating the US AI bubble? To weaken the stature of the USA on the global stage relative to themselves. Perhaps decrease US investment in AI and slow creation of some general AI superweapon I suppose. Because the goal is to show that cheap chinese AI can compete with expensive USA AI, it's nessisarially a low-cost attack relative to the "damage" it could create. >Why do we think there is…

  To weaken the stature of the USA on the global stage relative to themselves. Perhaps decrease US investment in AI and slow creation of some general AI superweapon I suppose.
I think you are overthinking it. The reason why Chinese AI labs have to go open source is because they do not have the clout to freely expand to international markets. Therefore, in order to succeed and get attention, they are providing their models for free if you have the inference hardware.

  Well that's the position that these chinese firms are trying to convince us of, and they can convince us by undercutting proprietary models in price/performance/openness.
I don't think these firms have said there is a US AI bubble and they're trying to pop it.

  Because they haven't bet the farm on AI like the USA has.
They are. The only problem is that they can't buy Nvidia chips or EUV machines so they're bottlenecked.

Re: The gap between open weights LLMs and closed source LLMs

#230

The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure mo…

“China can only copy the US” is a very short sighted and uninformed opinion. there is more coming out of china than just new ways to distill models

I don’t know how anyone can look at the innovation going on at DeepSeek and come to the conclusion that China can only copy.

Distillation and copying are how they’ve bootstrapped their models, but that feels not so different than Anthropic and Meta torrenting millions of pirated books.

The Chinese labs are solving problems for a different set of constraints.

Post reply on HN