Live data from Hacker News

Who's afraid of Chinese models?

stratechery.com

61–70 of 965 posts

Re: Who's afraid of Chinese models?

#61
post #2

> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service th…

The distillation explanation is classic American exceptionalism: No one could possibly do anything unless they were copying American leaders (where "American" means a bunch of Chinese, Canadian, Europeans and Indians working in the US). It's also a bit of securities defensiveness. Pretending that you really do have a super moat, people just keep swimming in it so you just need to add more alligators. It's farcical. A…

Which Chinese model was it that identified itself as Claude 15% of the time?

Re: Who's afraid of Chinese models?

#62

The people who are most afraid of Chinese models are the VCs who poured into Anthropic and OpenAI at astronomically high valuations. Anthropic is valued at $1.2T and OpenAI is targeting $850B. These astronomical valuations were built on the premise that these labs would generate massive profits from premium API pricing, but the Chinese labs are completely undercutting this strategy by releasing excellent open models…

The (quite excellent) article discusses several of your points. If you haven't read it, I recommend it.

  - Commodity market profitability is determined by marginal cost of production. LLMs have marginal cost; traditional software does not.

  - Models are not free. Downloading them is free. Running them is not. This has manufacturing economics, not software economics; the idea that they are "free" is an economic category error as it relates to their actual use

  - The highest tier Chinese models are not more economical than US frontier models. Try GLM 5.2 and see how much it costs to do real work. I did, and it was more expensive than GPT 5.6.

  - This is because US labs are leading on cost efficacy of inference ($/task)

  - Training will decline as a percentage of costs as inference expands compute share due to agentic workloads. A big part of training now is optimizing token efficiency. It's hard to distill token efficiency; that is perhaps why Chinese LLMs are so inefficient.

  - With increasing inference as % of total compute, if labs create efficient models -- which they can, because they can create highly optimized models amortized over very high inference loads -- they can be low cost producers, and be competitive at $/task rates
OpenAI really shows the way here. Their cost per task is less than half that of Anthropic because of more efficient tokenization and less verbosity. OpenAI is both cheaper and better than Chinese models for frontier work.

Re: Who's afraid of Chinese models?

#63
post #55

Earlier quoted context omitted.

Open weight models are much more auditable than closed models, but could still hide backdoors that could be near impossible to detect.

Correct. We need open weights, open code and open data. If nobody else can reproduce what someone did there will always be security questions. Even if we can reproduce it there could still be security concerns but it's more realistic to investigate yourself.

Exactly what are the possible 'security issues' of self hosting an open weights model?

Re: Who's afraid of Chinese models?

#64

The people who are most afraid of Chinese models are the VCs who poured into Anthropic and OpenAI at astronomically high valuations. Anthropic is valued at $1.2T and OpenAI is targeting $850B. These astronomical valuations were built on the premise that these labs would generate massive profits from premium API pricing, but the Chinese labs are completely undercutting this strategy by releasing excellent open models…

The (quite excellent) article discusses several of your points. If you haven't read it, I recommend it. - Commodity market profitability is determined by marginal cost of production. LLMs have marginal cost; traditional software does not. - Models are not free. Downloading them is free. Running them is not. This has manufacturing economics, not software economics; the idea that they are "free" is an economic category…

But this assumes Chinese models will not achieve token cost optimization. Intelligence needs are fairly flat for many tasks, and the Chinese models have caught up on this front. Next they achieve greater token cost efficiency and we don’t need OpenAI.

Re: Who's afraid of Chinese models?

#65

Earlier quoted context omitted.

The distillation explanation is classic American exceptionalism: No one could possibly do anything unless they were copying American leaders (where "American" means a bunch of Chinese, Canadian, Europeans and Indians working in the US). It's also a bit of securities defensiveness. Pretending that you really do have a super moat, people just keep swimming in it so you just need to add more alligators. It's farcical. A…

Which Chinese model was it that identified itself as Claude 15% of the time?

Models don't have some self identity, beyond what is explicitly handed to them via a system prompt. There have been many, many cases of models identifying as different models by different makers as a basic identity hallucination. They train on enormous volumes of data including lots of people talking about certain makers and models (ChatGPT was actually a super common one given that it became the kleenex of the LLM world). Hence why vendors have to specifically tell it to override that, and if they don't you get lots of funny cases of identity confusion.

This isn't the big gotcha some people seem to think it is, and the whole news cycle about that was mostly by people who have no idea what they're talking about. It's actually a meaningless data point. But it's precisely the sorts of people who think that a few thousand free accounts surreptitiously snuck off with Fable.

Re: Who's afraid of Chinese models?

#66
post #14

> It’s striking the extent to which Claude Code and Codex are proving to be quite sticky; whichever harness you start working with is likely to be the one you stick with, and that figures to be even more the case with non-technical users. My experience has been quite the opposite. I was using Claude Code almost exclusively this winter/spring and swapped to Codex earlier this summer. It took no time whatsoever to swit…

[dead]

Which would imply that these things are fast becoming… checks notes… a commodity?

Re: Who's afraid of Chinese models?

#67

The people who are most afraid of Chinese models are the VCs who poured into Anthropic and OpenAI at astronomically high valuations. Anthropic is valued at $1.2T and OpenAI is targeting $850B. These astronomical valuations were built on the premise that these labs would generate massive profits from premium API pricing, but the Chinese labs are completely undercutting this strategy by releasing excellent open models…

Good. Over the past few years, VCs have proven that they’re warmongering psychopaths. Hopefully China puts every last one of the Palantir/Flock/Anduril class out of business.

Re: Who's afraid of Chinese models?

#68
post #2

> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service th…

Forbidding distillation is like forbidding using a compiler to make another(perhaps better, more efficient) compiler.

Re: Who's afraid of Chinese models?

#69
post #60
post #42

Earlier quoted context omitted.

What kind of fresh hell does the sentence "undercut the professor" come from? Teaching isn't a race to the bottom. You don't undercut teaching by giving more lessons, just like you don't slight the hospital by performing CPR.

We are not really talking about teaching here.

No, we’re talking about an intimate set of tensors, not a human being.

A tree falling and killing someone isn’t tried for manslaughter.

So I don’t care about a hypothetical teacher.

Re: Who's afraid of Chinese models?

#70
> By the same token, don’t expect China to do anything about distillation attacks on the frontier labs. I think it is mistaken to attribute all of the success of Chinese labs to distillation, but it’s just as much of a mistake to pretend like distillation doesn’t give Chinese labs a big advantage.

I think we see this with Meta being paranoid about internal Claude usage, to avoid inadvertently distilling[1].

If distillation is a driver, then smaller American labs could be distilling, but are not for legal reasons.

But that's a big if we just don't know for sure.

1 - https://cryptobriefing.com/meta-restricts-claude-code-codex-...

Post reply on HN