Live data from Hacker News

“We have information that Moonshot distilled Fable for the development of K3”

twitter.com

691–700 of 742 posts

Re: “We have information that Moonshot distilled Fable for the development of K3”

#691

Earlier quoted context omitted.

> two wrongs don't make a right There is no second “wrong” here. Model outputs are not copyrightable. I think that was already established? Or do you think that Anthropic should own all the code generated by Claude? Surely that would be somewhat problematic? If Anthropic feels that some of their customers are breaking their EULA (nothing to do with copyright infringement though) they are free to stop doing business w…

there's a difference between a moral wrong and a legal wrong. in a heavily simplified view, moral wrongs are usually decided in the eyes of the victims -- you did bad thing to me so i'm not going to talk to you anymore. legal wrongs are decided by courts -- you did a bad thing so this court has decided you're not allowed to talk to that person anymore. anthropic are essentially saying in this tweet they believe a mor…

Sure but it’s hard to read what Anthropic is saying in any other way than that they think that it’s morally wrong to engage in any behavior that harms their (potential) profit margins. The exact phrasing is just a way to justify their stance to other people.

I mean you are right in a way of course, it’s just a matter of degree and perspective, though. If one thing is moderately morally wrong and the other is potentially lightly morally wrong I don’t think it’s fair to equate them.

To me the situation is a bit like Google coming out and saying that its morally wrong for someone to build a competing open operating system on top of Android while stripping all Google services and “stealing” their ad revenue. Just seems silly and hypocritical.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#692

Earlier quoted context omitted.

Generating a set of weights is not learning.

would you say that airplanes don't fly because they don't flap their wings? it's possible to achieve the same things with different approaches.

I would say airplanes fly, but I wouldn't say that submarines swim. Things have a bit more nuance, and the field of "learning" isn't as well understood as the ML proponents claim it is.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#693

Earlier quoted context omitted.

You're being too charitable to Anthropic, and assuming that the way they are abusing the word "distillation" has some real meaning here. It doesn't. Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training. You can't distill what you are not given - simple as that.…

> Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training. This is just straight-up factually false. The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, ma…

> I've personally repeatedly asked several different Chinese LLMs what their name is, and they answered "Claude"

I'm curious what you are doing to get them to override their own name that they were trained on and/or have as part of their system prompt?

I'd assume that the Chinese are scraping the internet for training data the same way western companies do, so for sure there will be a lot of AI generated content in their training data - you don't need to be paranoid and assume they must be getting it all direct from Anthropic.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#694

Earlier quoted context omitted.

Well there were observed cases of Claude calling itself Deepseek or Qwen. So pot calling the kettle black? To be fair I find it hard to take this too seriously, shouldn’t it be trivial to just replace “Claude” with any other string in your “distillation” dataset?

> Well there were observed cases of Claude calling itself Deepseek or Qwen. So pot calling the kettle black? There are open-source Deepseek and Qwen models - "distilling" doesn't involve breaking terms of service or hitting an API because you can literally run local inference or even just inspect the weights directly, and that's intended because they're open source . It's categorically different for a nation-state to…

> It's categorically

The fraud part and using stolen accounts or credit cards or blatantly violating the terms and conditions (i.e. reselling subscriptions not using outputs in certain ways somebody might not like) is indeed categorically different.

Using uncopyrightable outputs of an AI model obtained legitimately to train your model is not inherently interlinked with any of those things. I don’t really see how the model being proprietary or “open” is particularly relevant when talking about the outputs.

Even using the word “distilling” in this case is deceptive and biased. It implies that the Chinese are somehow stealing Anthropic’s models or their weights and somehow directly transforming them into new models. That’s certainly not what’s happening in any direct sense.

e.g. what if I agreed to send all my Claude code session files to Deepseek or whoever? There would be nothing wrong about that since I and not Anthropic own those files and can do whatever I want with them. Using certain different ways to obtain them of course could be highly illegal.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#695
post #570

Earlier quoted context omitted.

> It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ... > ... but Chinese SOTA foundries directly using distillation as fair game. As someone who says it’s fair game, it’s less that I’m being hypocritical and more that I don’t care that one thief had their shit stolen by a second thief. I also wouldn’t ca…

Why are you calling the Chinese companies “thieves” though? LLM outputs are not copyrightable (or rather the user is effectively the only one who can own it). It would be problematic if Anthropic owned all the software generated using Claude..

They are thieves the same way Anthropic or OpenAI are thieves. Either it’s fair use to learn from this data or not.

If what Anthropic/OpenAi did for training is theft then the Chinese models also are a form of theft. If they didn’t steal then I don’t think the Chinese firms did either.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#696
post #581

Earlier quoted context omitted.

And some people might want to know about [insert your favourite beyond-the-pale-in-the-US topic here, we're on a US forum after all]. I think it would be great if those people could turn to Chinese models, while anyone wanting to know about Uyghur camps can ask the US ones.

> beyond-the-pale-in-the-US topic Can you name one? It's an honest question, first of I'm not American, but also the stuff I do come up with (asking an LLM how to blow up a school or whatever) would also be handled similarly in China, so those would be a wash, and I can't think of any that aren't.

I would imagine a lot of things touching upon progressive politics would be affected (there were a few high-profile incidents demonstrating bias like Google's black Wehrmacht soldier pictures, but has anyone rigorously tabulated how the various commercial LLMs respond to questions about the gender binary or heritability of human traits considered good or bad?). Overall, I'm too reluctant to even write out in the abstract sequences of words that I never want to explain to a future job interviewer or HR employee who used GPT-7 to comb the internet for all text that stylistically can be traced to me, but just imagine whatever you believe to be vile and wrongheaded opinions in that general space which it is certainly more than justifiable to prohibit. The things that make you think "banning this is good actually" are exactly the things most likely to be banned (and this is true in China too).

Another thing I would try if I had access to the models and enough proxies to hide behind is asking for advice on software/movie piracy or seeing to what extent the models can be elicited to straight up argue against the validity of intellectual property, though there it seems more probable to me that the US models would be permissive.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#697
post #137

Earlier quoted context omitted.

> on the other hand the complete dismissal of copyright by AI labs Courts keep ruling over and over that an LLM trained on copyrighted works qualifies as a transformative work and is therefore fair use. They don't have to dismiss copyright law, this has always been allowed. The only thing they get in trouble for is pirating the works to get their hands on them.

Man that's depressing to read someone defending this

I'm telling you what the law plainly says. I personally think that the exceptions for fair use made sense before the transformer was invented, and they still make sense now.

The two main problems with copyright also haven't changed: copyright lasts too long and is too expensive to defend.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#698

Earlier quoted context omitted.

Apple and Samsung are ~40-50% of the global market, depending on the source. Saying that smartphones are mostly from Chinese companies isn't accurate. And Samsung and Apple are gaining market share, while the major Chinese brands are losing market share. https://www.idc.com/promo/smartphone-market-share/ https://gs.statcounter.com/vendor-market-share/mobile/worldw...

iPhones are predominantly made in China Foxconn facilities in cities like Zhengzhou (known as "iPhone City") and Shenzhen.

Yes, but I was responding to this comment:

"Smartphones are mostly Chinese now (except for iPhones and Samsung)."

Which implies that we are talking about the companies, not where they are manufactured.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#699
post #514

Earlier quoted context omitted.

You're both reading tea leaves. Settling just says that they expected the internal costs or risks to be more than 1.5 billion cashflow. In the $65B in Series H funding at $965B post-money valuation they said their run-rate revenue crossed $47B annualised. With those numbers, there can be sound financial reasons for wanting to just get rid of the lawsuit. Also if it ends up that other competitors also need to pay $1.5…

> With those numbers, there can be sound financial reasons for wanting to just get rid of the lawsuit. Yes, of course, on Anthropic's side. Why would the other side agree to a settlement?

Because most authors don't make a ton of money from their books, and even a relatively low settlement from Anthropic is a large enough sum that they're OK with taking it.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#700
post #675

Earlier quoted context omitted.

Encrypted reasoning traces don't prevent distillation. You only need the input prompts and final output responses to distill capabilities. As Anthropic explained in their distillation report (linked above), Moonshot AI already collected millions of session traces, covering: * Agentic reasoning and tool use * Coding and data analysis * Computer-use agent development * Computer vision Hiding the internal CoT blocks sto…

How would you propose to distill Fable-style input/output pairs without the CoT? If you use them as SFT input, you’ll be trying to train a model to predict the post-reasoning output without any reasoning, and this seem very unlikely to work at all with the size of model that Kimi produced and the complexity of Fable’s output. You can’t really “RL” with them because they would be so far off policy that there would be…

Anthropic removed raw CoT from all models since Sonnet 3.7 (released February 2025), all Opus 4.X models have censored CoT, yet Chinese labs continue to distill from them, proving this restriction wasn't an insurmountable hurdle.

There are some outputs that are highly valuable and cannot be censored, such as tool API calls or agentic tool usage, which is precisely what Moonshot was accused of harvesting/distilling.

Tool call harvesting was enough of an issue that Anthropic began inserting spurious tool calls when they detected distillation attempts, trying to poison the distilled data.

There is certainly more data the labs are training off of, but that's difficult to know without insider information.

Post reply on HN