Live data from Hacker News

Who's afraid of Chinese models?

stratechery.com

41–50 of 965 posts

Re: Who's afraid of Chinese models?

#41
post #2

> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service th…

The distillation explanation is classic American exceptionalism: No one could possibly do anything unless they were copying American leaders (where "American" means a bunch of Chinese, Canadian, Europeans and Indians working in the US).

It's also a bit of securities defensiveness. Pretending that you really do have a super moat, people just keep swimming in it so you just need to add more alligators.

It's farcical. Anyone who has worked on large models knows that the premise that an almost-Fable model was trained with distillation is beyond ridiculous. It's theoretically possible if they spent tens of billions of dollars on API calls, but it isn't the magic that somehow these people keep convincing people it is.

Previously Anthropic has reported on some Chinese firms doing chicken-shit level of API calls, that at most would be doing some Q and A or final fine tuning. The notion that they're training these models via it is fantastically ignorant nonsense that only very ill-informed and gullible people fall for.

Re: Who's afraid of Chinese models?

#42
post #34

Earlier quoted context omitted.

This is a silly perspective, inaccurate, and out of bounds framing. Public libraries, in this instance, is curated data from all the internet, obtained through not legal means (I don't have a problem with this other than lack of attribution, being copy-left). Just to be clear. But in answer to your incredibly leading and inaccurate framing... they are required (by their job title) to teach to those who who show up in…

If a professor learns from multiple books, generalizes from them and then shares his knowledge he is providing a valuable service. Versus someone who makes a recording of the professor's lectures and resells them to undercut the professor--that guy is not providing a valuable service.

What kind of fresh hell does the sentence "undercut the professor" come from?

Teaching isn't a race to the bottom. You don't undercut teaching by giving more lessons, just like you don't slight the hospital by performing CPR.

Re: Who's afraid of Chinese models?

#43
post #34

Earlier quoted context omitted.

If a professor learns from multiple books, generalizes from them and then shares his knowledge he is providing a valuable service. Versus someone who makes a recording of the professor's lectures and resells them to undercut the professor--that guy is not providing a valuable service.

What about a student attending lectures of other professors and generalizing what he learns from them, then going on to become a professor?

If the student is really good at generalizing we wouldn't even be having this debate because he would've just generalized from the same source materials the professor used.

Re: Who's afraid of Chinese models?

#44

According to openAI's own @deanwball: Even OpenAI isn't buying this distillation talk: https://xcancel.com/deanwball/status/2078133895766114412#m

Can you or someone please explain several of the claims made in this tweet? "I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks" what risks? I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). Confused what this means Open-weight models are inherently dec…

> I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). Confused what this means

I think it's referring to the belief that LLMs are not the path towards AGI, and that LLM's, while useful, are not going to have the impact that the American labs believe it will have.

Re: Who's afraid of Chinese models?

#45
post #26
post #2

> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service th…

Seems only fair that if LLMs can use copyrighted data for training then they should be able to use cannot-be-copyrighted output of other LLMs. But barring the terms of service from forbidding distillation seems like a tough sell. OpenAI shouldn't be allowed to decide what types of customers it wants and doesn't want?

This happens all the time. The government can decide legislatively that certain commercial terms are simply unenforceable. Making distillation clauses unenforceable in tort law would be straightforward. They can decide what customers they want to have, but they do not have unfettered rights as to the enforceability of terms governing the relationships between the parties.

Re: Who's afraid of Chinese models?

#46
post #14

> It’s striking the extent to which Claude Code and Codex are proving to be quite sticky; whichever harness you start working with is likely to be the one you stick with, and that figures to be even more the case with non-technical users. My experience has been quite the opposite. I was using Claude Code almost exclusively this winter/spring and swapped to Codex earlier this summer. It took no time whatsoever to swit…

I think they stickiness is less about the difficulty of switching and more about the lack of desire. I’ve been using Claude since day one, it works well and I’m happy, I like it. I’m sure Codex is good too. Switching from one to the other certainly isn’t going to be a game changer, the discourse shows me the differences are marginal. Probably the only reasons I would seek change are economical.

Convergence in coding makes them highly substitutable. But I could see harnesses configured for different purposes -- let's say, a harness for creating teaching plans -- being able to cater to its audience better than a coding harness. Maybe it's got tools to plug into standardized curricula, what the lesson books will be, what other lesson plans the district's teachers have made, etc., which could be done in a clunky way in a regular harness but could be streamlined.

Re: Who's afraid of Chinese models?

#47
post #13
post #2

> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service th…

[flagged]

More like, is a professor who learned from books prohibited from writing his own books on the subject?

Re: Who's afraid of Chinese models?

#48
post #19

One thing I have not seen mentioned between Chinese AI vs US, population. China has a billion+ people that their AI can "study". Plus due to China's political structure, their AI has access to everyone's chats, comments and sites, scraping everyting. Here in the US, with 1/3 the population, the AI race was lost before it even began. Plus in the US, all companies and people are doing all they can to restrict AI from s…

[deleted]

Re: Who's afraid of Chinese models?

#49
post #34

Earlier quoted context omitted.

This is a silly perspective, inaccurate, and out of bounds framing. Public libraries, in this instance, is curated data from all the internet, obtained through not legal means (I don't have a problem with this other than lack of attribution, being copy-left). Just to be clear. But in answer to your incredibly leading and inaccurate framing... they are required (by their job title) to teach to those who who show up in…

If a professor learns from multiple books, generalizes from them and then shares his knowledge he is providing a valuable service. Versus someone who makes a recording of the professor's lectures and resells them to undercut the professor--that guy is not providing a valuable service.

What are you even trying to say: "undercut the professor"...

The further you try to constrain this topic into this illformed analogy the weirder it becomes. If we start off with a better analogy...

Re: Who's afraid of Chinese models?

#50
post #13
post #2

> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service th…

[flagged]

1) No one is asking Anthropic to give tokens for free, but at market rates.

2) Any professor who tried to ban students from posting lecture notes online would be immediately mocked.

Post reply on HN