Live data from Hacker News

China’s open-weights AI strategy is winning

werd.io

271–280 of 978 posts

Re: China’s open-weights AI strategy is winning

#271

I do think open-weights models are going to "win" in the sense that they're probably going to be dominant when the hardware to run them becomes affordable. (which might be a while). Although I guess you could probably rent the GPU's yourself to hypothetically save on costs. (I'm a little skeptical -- I've heard of companies doing this and the inference bills are surprisingly high -- assuming the sources are correct.…

> I just don't really understand the business model behind it

There’s a lot of value in the same sense there is a lot of value in controlling what Google search results are shown and what people see in the Twitter feed.

Re: China’s open-weights AI strategy is winning

#272

I do think open-weights models are going to "win" in the sense that they're probably going to be dominant when the hardware to run them becomes affordable. (which might be a while). Although I guess you could probably rent the GPU's yourself to hypothetically save on costs. (I'm a little skeptical -- I've heard of companies doing this and the inference bills are surprisingly high -- assuming the sources are correct.…

They get money from subscriptions and tokens, same as for closed-weight providers. Yes they'll lose some traffic to hosting services, but many users prefer to use the original training company since they have a guaranteed-correct implementation. Similar business model as open-source SaaS companies.

Some companies (most notably Deepseek) also manage to host their own LLMs so efficiently they undercut all third-party hosting services.

Re: China’s open-weights AI strategy is winning

#273

I do think open-weights models are going to "win" in the sense that they're probably going to be dominant when the hardware to run them becomes affordable. (which might be a while). Although I guess you could probably rent the GPU's yourself to hypothetically save on costs. (I'm a little skeptical -- I've heard of companies doing this and the inference bills are surprisingly high -- assuming the sources are correct.…

I think part of it is definitely to weaken US providers and the US economy as a whole. China has a completely different domestic economic structure and motivations from western cultures... it's probably closest to a fascist economy mixed with a Maoist cultural ideology behind it. There's definitely winners and losers and the state tends to have tight controls over everything though.

I also think the restrictions on OpenAI and Anthropic are somewhat short sighted. In that the guardrails dramatically limit efforts towards securing your own software in many ways. Yes, it's also "dangerous" and maybe there should be a means of identifying "domestic" or otherwise "secure" accounts for those allowed to use the models without the same guardrails in place.

Re: China’s open-weights AI strategy is winning

#274
post #6

I’m suspicious of some quotes here, “80% of startups using Chinese models,” doesn’t seem quite right to me. I just interviewed at several startups and they were all using the US models. Maybe they have some minor use of Chinese models but the bread-and-butter of most of these businesses model use is the Claude and Codex subscriptions.

Yeah. People should absolutely be _trying_ the Chinese models, and experimenting with running things locally, but the noise in development is genuinely all Claude and Codex. I put my foot in the mobile comparison the other day, and will again. If you were to go back and be a mobile dev in 2010 by all means specialize on one platform, but play with both as a professional interest to stay realistic. Here it's important…

Unlike iOS/Android choice which has broad personal ecosystem implications, changing IP address to another LLM provider is effortless.

Re: China’s open-weights AI strategy is winning

#275
post #6

I’m suspicious of some quotes here, “80% of startups using Chinese models,” doesn’t seem quite right to me. I just interviewed at several startups and they were all using the US models. Maybe they have some minor use of Chinese models but the bread-and-butter of most of these businesses model use is the Claude and Codex subscriptions.

Yeah. People should absolutely be _trying_ the Chinese models, and experimenting with running things locally, but the noise in development is genuinely all Claude and Codex. I put my foot in the mobile comparison the other day, and will again. If you were to go back and be a mobile dev in 2010 by all means specialize on one platform, but play with both as a professional interest to stay realistic. Here it's important…

Dunno I put to deep seek the question what I’d need in hardware to get full service k3 with lower latency in western pa (I was making like it was a business proposition.) It says several million dollars at minimum.

Re: China’s open-weights AI strategy is winning

#276

AI models cost tens of millions to train. Offering them for free won’t justify the upfront costs. The Chinese model of model training/open sourcing only makes sense in the context of the overall strategy of undercutting American frontier labs’ profit margins.

There is a huge cultural influence opportunity too. Imagine if, in 10 years time, every school kid is learning the causes of the US civil war from an LLM, getting their essays on hiroshima and nagasaki graded by an LLM, and a million other things. A country with competitive LLMs gets to decide whether "it was more complicated than just slavery", and whether "it was tragic but necessary, saving lives over all". Countr…

Maybe your kids. In my corner of the country, parents are actively demanding that schools back off computer usage, much less AI

Re: China’s open-weights AI strategy is winning

#277

This is a very strange article considering that Llama, the mother of all open-weight models, has led to anything but success for Meta. Also, enterprises don't give a rip if models are open. They care about zero data retention (and sticking with whatever vendor they're already using). This blog post is suspiciously close to being a restatement of what Alex Karp recently said on CNBC[0]. It's important to remember he's…

The behind the scenes element you're not aware of is Anthropic is going to companies reliant on their models and demanding HUGE one time fees (100 million+) to continue using their models or they will be cut off. This has happened to several larger companies I and others are invested in. This resulted almost every time in "screw off we'll train our own models or use refined open source ones instead" leading to a lot…

I am not up-to-date in this area, and not necessarily that I don't trust you, but do you have a source for this? Just curious.

Re: China’s open-weights AI strategy is winning

#278

I do think open-weights models are going to "win" in the sense that they're probably going to be dominant when the hardware to run them becomes affordable. (which might be a while). Although I guess you could probably rent the GPU's yourself to hypothetically save on costs. (I'm a little skeptical -- I've heard of companies doing this and the inference bills are surprisingly high -- assuming the sources are correct.…

> I'm sort of baffled by what the entities that train the open-weights models get out of it though. Is it just a direct play to undercut the US providers because they view them as a threat? I just don't really understand the business model behind it.

In China, it's because they are being heavily subsidized to do the research activity. It's not really complicated -- if you allocate public money for people do to a thing, they will do it.

Re: China’s open-weights AI strategy is winning

#279

The US business model for commercializing LLMs seems unsustainable to me. We are saying that they are creating trillions of dollars in value out of: 1. A model that for the most part is public and available to anyone. 2. A situation where the model’s success mostly comes from throwing as much data and computational resources at it as possible. It seems that either of those assumptions could crumble quickly and unexpe…

There's already cases where Google and I'd assume others are designing chips to work with specific models more efficiently in coordination... Personally, I could even see more specialty models called/coordinated from the larger models that can do smaller pieces of targeted work very well within limited scopes as a mixed economy so to speak.

Assembling a new model from scratch requires a ton of resources and knowledge bases... there's been a lot of sketchy activity just in training. You also have weighting, distillation and other approaches to create more portable options that can run on lesser hardware. But, K3 as an example takes massive compute resources to run.. and this isn't going to get to a portable device any time soon... as Moore's law is effectively dead, you may get newer/better tooling around the LLMs, or you may get an entirely new/unique approach to AI... but current trends aren't going to put a leading model on your own hardware anytime soon for most people.

Re: China’s open-weights AI strategy is winning

#280

Earlier quoted context omitted.

How much would it cost (time and resources) to take a Chinese open-weight model and remove these (admittedly) stupid guardrails?

You can find on Huggingface a huge number of Chinese open weights LLMs from which the censorship has been removed. They typically contain in their names words like -abliterated or -uncensored. For some of the recent bigger Chinese LLMs, it took a longer time until someone succeeded to remove the censorship, but eventually uncensored variants were published. E.g. for Kimi 2.6 an uncensored variant appeared only a coup…

It was a rhetorical question. OP is making it sound like the open weight models are fundamentally broken by being censored out of the box. This is a completely asinine take.
Post reply on HN