Live data from Hacker News

Who's afraid of Chinese models?

stratechery.com

51–60 of 965 posts

Re: Who's afraid of Chinese models?

#51

People seem to conflate "made in China" with "can't be trusted." id argue the bigger distinction is open vs. closed. An open model can be audited, fine-tuned, and technically run entirely on your own hardware. A closed model is basically "trust us."

Open weight models are much more auditable than closed models, but could still hide backdoors that could be near impossible to detect.

Re: Who's afraid of Chinese models?

#52
I operate an analytics site (pretty big one B2B where client's backend feeds data into our system), and we see tons of traffic originating from northwestern China (Xinjiang) from Shenzhen Tencent Computer Systems Company Limited.

There are also half a dozen other companies from China continuously hammering our clients’ websites.

I was wondering, what's in that cold dessert? Low and behold satellite imaging shows massive datacenter build outs, very cheap solar energy.

Few months ago something happened and the Geo location on data on those IP now shows "Shanghai" or "Shenzhen". A way to cover tracks? But mapping latency still points to fact that nodes behind these IPs are still operating around Xinjaing region

credit:

'You Can't Cheat Time: Finding foes and yourself with latency trilateration' https://youtu.be/_iAffzWxexA HN user: lopoc

Shenzhen vs Xinxiang is hard to do using this technique but Shanghai vs Xinxiang does show difference.

Assuming that China only distills is a huge mistake.

It’s no longer some backward place that does low value copying. Look at companies like ByteDance and Xiaomi.

Chinese companies aren’t just distilling, they’re acquiring data in the same way American companies did by paying people and crawling the internet.

The way I understand it, China has a few large companies that crawl the web at a rapid rate and build corpora. The government essentially wants select few companies to do this and then make the data available to other strategic companies operating within China.

Then there are data aggregators that buy data from apps, websites, and services, as well as systems like OpenRouter or Cursor, where companies can learn from the “traces” of coding agents, chats, and so on.

This massively reduces costs, as smaller companies like DeepSeek don’t have to do their own crawling or acquire data from 100s of websites and coding agents etc....

There are also companies in China that buy American LLM APIs and proxy them to companies within China. So, there could be 10,000+ companies using American AI products, while China logs all of this, understands how they’re being used, and trains on their traces.

Re: Who's afraid of Chinese models?

#53

The people who are most afraid of Chinese models are the VCs who poured into Anthropic and OpenAI at astronomically high valuations. Anthropic is valued at $1.2T and OpenAI is targeting $850B. These astronomical valuations were built on the premise that these labs would generate massive profits from premium API pricing, but the Chinese labs are completely undercutting this strategy by releasing excellent open models…

The valuations are unjustified even at the prices they’re charging now. They’re going to try their best to offload these investments into our pensions before the inevitable crash.

Right, but retail investors weren’t supposed to find that out until after the IPO.

Re: Who's afraid of Chinese models?

#54
post #2

> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service th…

Government cannot exactly "bar" terms of service. ToS isn't law. The most they can do is say they're unwilling to enforce them.

ToS is just conditions that you agree to in order to use a private service that is provided at-will. I can have a private coffee shop where the terms of service are that you must wear red to enter, and if you're not wearing red, you are not welcome on my property.

So it would be upto OpenAI and Anthropic to enforce them on their own terms (by banning accounts and IPs).

Re: Who's afraid of Chinese models?

#55

People seem to conflate "made in China" with "can't be trusted." id argue the bigger distinction is open vs. closed. An open model can be audited, fine-tuned, and technically run entirely on your own hardware. A closed model is basically "trust us."

Open weight models are much more auditable than closed models, but could still hide backdoors that could be near impossible to detect.

Correct. We need open weights, open code and open data. If nobody else can reproduce what someone did there will always be security questions. Even if we can reproduce it there could still be security concerns but it's more realistic to investigate yourself.

Re: Who's afraid of Chinese models?

#56
post #14

> It’s striking the extent to which Claude Code and Codex are proving to be quite sticky; whichever harness you start working with is likely to be the one you stick with, and that figures to be even more the case with non-technical users. My experience has been quite the opposite. I was using Claude Code almost exclusively this winter/spring and swapped to Codex earlier this summer. It took no time whatsoever to swit…

Same here. I flip flop between them. Most people I know who have access to both, technical or not, are doing the same. They’re just too close and sometimes one does what you want better than the other.

Re: Who's afraid of Chinese models?

#58
post #13

Earlier quoted context omitted.

[flagged]

More like, is a professor who learned from books prohibited from writing his own books on the subject?

He is prohibited from regurgitating source material, of course! But if he generalized from the books he read and really learned the subject--and even made new connections between ideas--then he is free to write his own book.

Re: Who's afraid of Chinese models?

#59
post #54
post #2

> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service th…

Government cannot exactly "bar" terms of service. ToS isn't law. The most they can do is say they're unwilling to enforce them. ToS is just conditions that you agree to in order to use a private service that is provided at-will. I can have a private coffee shop where the terms of service are that you must wear red to enter, and if you're not wearing red, you are not welcome on my property. So it would be upto OpenAI…

The government absolutely can pass laws that ban particular contract previsions. They do that all the time. In your analogy for example while they can require you to wear red, they can't require you to be white.

Re: Who's afraid of Chinese models?

#60
post #42
post #34

Earlier quoted context omitted.

If a professor learns from multiple books, generalizes from them and then shares his knowledge he is providing a valuable service. Versus someone who makes a recording of the professor's lectures and resells them to undercut the professor--that guy is not providing a valuable service.

What kind of fresh hell does the sentence "undercut the professor" come from? Teaching isn't a race to the bottom. You don't undercut teaching by giving more lessons, just like you don't slight the hospital by performing CPR.

We are not really talking about teaching here.
Post reply on HN