Live data from Hacker News

The gap between open weights LLMs and closed source LLMs

blog.doubleword.ai

111–120 of 264 posts

Re: The gap between open weights LLMs and closed source LLMs

#111

The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure mo…

Coding a case where it's possible to programmatically generate large amounts of data relatively cheaply. China could realistically surpass the US in coding while still being behind in many other areas.

Also worth noting that China has more data to work with in general having a much bigger population.

Re: The gap between open weights LLMs and closed source LLMs

#112
post #81

I haven’t seen it discussed anywhere that closed models can essentially cheat benchmarks right? What Anthropic or OpenAI brand as a model doesn’t necessarily have to be just weights, it can be a whole backend system that augments the model itself. With this they can score better benchmarks than an open source model that is weights alone.

Good point

Re: The gap between open weights LLMs and closed source LLMs

#113

Earlier quoted context omitted.

> they can never be taken away Your right to 3d print whatever you want is about to be taken away (in California). What software you can run on your computer can already be restricted. Absolutely everything can be taken away. The simplest way to remove open models is probably to declare them a tool that terrorists could use. Crazy? Yes, the world is totally crazy these days.

Just like declaring piracy illegal stopped piracy and removed pirated materials from everyone's computers. Everything cannot, in fact, be taken away. Don't propagandize yourself. Some things, like information, are free. Not even China can prevent all its citizens from accessing Western internet. USGov simply does not have the resources to find and audit every hard drive and USB stick in the country for illegal files.…

You wouldn't download a car?

Re: The gap between open weights LLMs and closed source LLMs

#114
post #80
post #14

Interesting to consider this inline with recent us export bans, could the US be squandering its lead by giving the open source, largely Chinese labs catch up (in terms of model quality available to masses), will US labs be able to maintain the lead without users being able to use their latest models?

Why do you think this matters? Not that it does or doesn't but what quality does "US WINS" or "CHINA WINS" bring to the table?

I think the unspoken fear is that if we assume one or the other will "win" in reaching AGI(or whatever threshold of capability), the rest of the world will sooner or later live under their system of rule as a consequence

Re: The gap between open weights LLMs and closed source LLMs

#115

The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure mo…

This seems wildly naive. This entire field is like 4 years old. We have quite frankly no idea about what things will look like in 4 more years.

The article makes a very specific claim with a clear deadline less than 6 months ago. I do not underestimate the Chinese labs and their capabilities, if they wish they can retool to start overtaking the US labs with a different strategy. My comment shouldn’t be read as a permanent impossibility statement, just an observation on where we are right now. At the moment their strategy seems to be to produce decent quality, highly optimized models; and a pivot will take longer than 6 months to materialize into overtaking the frontier labs (that themselves do not look like they will throw the towel in in the next 6 months).

Re: The gap between open weights LLMs and closed source LLMs

#116

Earlier quoted context omitted.

It's just a smart business decision that allows their models to compete and gain market-share against much pricier private models. No philanthropy there.

It depends how you define philanthropy - obviously corporations don't just donate such valuable products to the world to make it a better place, but in effect that's what they end up doing in their effort to gain market share or brand recognition. Actual human philanthropists are sometimes doing it for the similar reasons of self-promotion.

Open source, Open weights, these are core business decisions.

Re: The gap between open weights LLMs and closed source LLMs

#117
post #78

The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure mo…

The amount of data Anthropic has claimed was extracted for distillation is tiny in comparison to the entire internet, which is right there for the taking and holds most of the knowledge people expect models to have. Distilling even with small amounts of data from a better model is still helpful, but not in the sense of transferring capabilities the raw internet-trained model doesn't have at all, but for identifying t…

I think you are making a distinction between pre training and later stages? The value on eg Fable output is exactly the careful preference optimization embedded in those responses. Not all data is the same (sorry if my first comment was sloppy on that).

Re: The gap between open weights LLMs and closed source LLMs

#118
post #81

I haven’t seen it discussed anywhere that closed models can essentially cheat benchmarks right? What Anthropic or OpenAI brand as a model doesn’t necessarily have to be just weights, it can be a whole backend system that augments the model itself. With this they can score better benchmarks than an open source model that is weights alone.

Sure, I think that's fine, that all counts. It counts for open source too, it's not like they're somehow running these benchmarks without any harness.

Nobody cares if your AGI is 100% made out of neural networks or if it's like 50% neural networks and 50% perl scripts.

Re: The gap between open weights LLMs and closed source LLMs

#119
post #47

The Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure mo…

> Chinese labs must entirely retool from harvesting frontier model data to producing the data systems and efforts to produce novel data Even if your characterization is accurate, they could do this tomorrow and are not so myopic that they wouldn’t have thought about it. I don’t see this as a barrier, and I see a lot of the same underestimation of Asia that’s been happening for 50 years. There’s not some innate Americ…

I am not sure which part you are interpreting as underestimation or whatever? Quite the opposite: I claim the difference arises from a difference in strategies, not from intrinsic differences in ability.

Also I was responding to a claim about what will happen in less than 6 months (that’s about the edge of what you can meaningfully say too much about in this field).

These strategies take materially different resources; it’s not an overnight decision made by leadership. I suppose there is a natural experiment ongoing at Meta regarding this, it seems they recently moved a number of people into a division to produce such data overnight. So we will find out soon how quick they climb the leaderboards.

Re: The gap between open weights LLMs and closed source LLMs

#120
I just hope CCP doesn't follow the US government and won't pull the plug before their companies release something on-par with the US frontier models. The question is whether US models not available to the general public will count.

The question is not whether they'll prohibit open-weight models better than the US ones, because we all know the obvious answer.

Post reply on HN