Live data from Hacker News

China’s open-weights AI strategy is winning

werd.io

471–480 of 978 posts

Re: China’s open-weights AI strategy is winning

#471

Earlier quoted context omitted.

> 15 months ago, the top model on the Artificial Analysis index was GPT-o3. It scores 30 on the Artificial Analysis index. There must be something really of with those benchmarks. Yes, hallucinations gotten better, but I don't see that the big frontier models got so much better in the last 12-18 Months. They just put out bigger wall of texts and feel smarter. But they still make way too many stupid errors

12 months ago "way too many stupid errors" was constant news. Today, you rarely hear about those anymore. Sure, the novelty of the errors has worn off a bit and thus the reporting. Nevertheless the quality has improved immensely in this regard. Also, AI video generation is now so good and accessible that it is very, very regularly used for memes, disinformation and proper (short) movie projects. AI image generation e…

Maybe it got a lot less and I just got used to it. True.

Still feels too much for me. Breaks my workflow for no reason. Too much overhead for me, if I can't trust the output

Re: China’s open-weights AI strategy is winning

#472
post #450

Earlier quoted context omitted.

You could say that both of them stole, but different stuff.

You can't steal intellectual property, only infringe on the copyright holder.

The Chinese models poked and prodded better models for training data and to avoid having to pay humans to RLHF themselves. You can call it infringement, theft, whatever, but it’s quite obviously “not ethical” to me.

Re: China’s open-weights AI strategy is winning

#473

The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins. - PCs destroyed minicomputers. Mainframes survive, but serving a much tinier portion of the market than they used to. - PC office productivity software destroyed expensive professional products. - Windows (low end) and Linux (free) completely destroyed the UNIX marketplace, and again, have taken huge mar…

The value of an LLM is the dynamic reasoning you get out of it and the cost to execute on that. I see two forces working against this that proprietary models will always have over an open source model. 1. The biggest is content licensing. Content is quickly becoming gated by systems at the front of their load balancers, completely changing the social contract of the Internet. What used to be a quick google search for…

> A 3rd party hosting platform like fireworks.ai that hosts open weights models won't provide any of that at all.

Why couldn't they?

Re: China’s open-weights AI strategy is winning

#474

Earlier quoted context omitted.

> It’s in the same neighborhood but isn’t really apples to apples. Distilling LLMs is to take a synthesized result that comes from huge amounts of innovation and computation, while the other is scraping what already exists as is. Hang on, why is scraping the public pool of knowledge not taking "a synthesized result that comes from huge amounts of innovation and computation"? You think that that all those github repos…

The Chinese models are the result of just as many papers, GitHub repos, etc… AND the synthesized results of those.

> The Chinese models are the result of just as many papers, GitHub repos, etc… AND the synthesized results of those.

Right, but they aren't the ones whining that other people are getting "the synthesised results" for free.

Re: China’s open-weights AI strategy is winning

#475
post #386

Earlier quoted context omitted.

The problem (right now) is that Open Weight models depend right now on huge companies to spend billion of dollars to train and develop them, all backed up by their incentives and their state to support this, while essentially giving away their monetization path. With open source projects, the benefit was that each individual could improve the complex system (e.g. Linux Kernel) interpedently, and over time the benefit…

> The problem (right now) is that Open Weight models depend right now on huge companies to spend billion of dollars to train and develop them, all backed up by their incentives and their state to support this, while essentially giving away their monetization path. Right... and there are two problems with this: 1. Eventually the capabilities of closed-weight models will just vastly outstrip open-weight models if the u…

> Open-source software isn't a moral good

Yes it is.

Re: China’s open-weights AI strategy is winning

#476

Earlier quoted context omitted.

You could say that both of them stole, but different stuff.

Don’t forget that the Chinese models are also built on top of huge amounts of “stolen data” as well, beyond the distilled. So it’s basically all of the above. However, there’s no mechanism for the NYT or an author or anyone in the US to sue the Chinese companies that took their work.

and if you're an author that lives outside the United States?

Re: China’s open-weights AI strategy is winning

#477
post #428

Earlier quoted context omitted.

Try instructing Codex to (say) fine-tune a language model based on a collection of books you've got saved. You will find yourself admonished, repeatedly and at length, not to utilize copyrighted materials to train language models, by an AI who owes its entire existence to that very act. These models might be smart but they're not close to being able to savor irony.

This behavior is actually specific to ChatGPT because they lost a music copyright lawsuit in Germany. They would refuse to output music lyrics too but they would happily do analysis on lyrics if you supply them. I suspect there might be a guardrail model involved here.

The trouble is, even if they refuse to output that copyrighted material, they were still trained on it without proper licensing and will still produce derivative work based on them because that's how this whole thing works.

Re: China’s open-weights AI strategy is winning

#478

The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins. - PCs destroyed minicomputers. Mainframes survive, but serving a much tinier portion of the market than they used to. - PC office productivity software destroyed expensive professional products. - Windows (low end) and Linux (free) completely destroyed the UNIX marketplace, and again, have taken huge mar…

if you look at how GPU memory grew in the last 15 years, it's about 10x. Sadly, 10x from today doesn't get us to a typical frontier model size of today which is a quickly moving target. some other advancement needs to happen to get us another 10x both in memory/compute requirements, and also power requirements.

Capability per GB and per watt has also been going up lot. This will continue in the future as well (not necessary as the same rate as last years). But enough that I think Opus 4.8 level is reachable on consumer PCs within 10 years from its release. Say at the price point of 2000 USD in 2025 dollars.

Re: China’s open-weights AI strategy is winning

#479

The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins. - PCs destroyed minicomputers. Mainframes survive, but serving a much tinier portion of the market than they used to. - PC office productivity software destroyed expensive professional products. - Windows (low end) and Linux (free) completely destroyed the UNIX marketplace, and again, have taken huge mar…

> free and low-end eventually wins

Not in SaaS which is what LLMs are. You can get VMs for much cheaper than AWS, Microsoft, and Google offer them but large companies (and startups) are happy to pay a premium for the support, reputation, and reliability that they perceive those companies as offering. Same thing for some of the managed database providers who are effectively selling a very heavily marked up version of postgres.

> The high price, and social pushback, mean that the American companies producing these models are precarious

I doubt it. The models really aren't that expensive when you look at what they can do. Fable is probably at least as good as the average software engineer and costs $50/wk on the max plan vs a software engineer who would cost closer to $4000 a week. The real money is probably in selling to enterprise vs consumers (Google has best route to making money from consumers since they can do what they did with ads and search to LLM queries).

It seems unlikely to me that US companies will send important corporate data to models controlled by a Chinese company as well.

Re: China’s open-weights AI strategy is winning

#480
post #451

Earlier quoted context omitted.

Try instructing Codex to (say) fine-tune a language model based on a collection of books you've got saved. You will find yourself admonished, repeatedly and at length, not to utilize copyrighted materials to train language models, by an AI who owes its entire existence to that very act. These models might be smart but they're not close to being able to savor irony.

I was a little radicalized when ChatGPT literally refused to translate parts of 1000+ year old religious texts and told me it was due to copyright concerns.

Cant you in this case point out that obviously it is an old twxt and there is no copyright?
Post reply on HN