Live data from Hacker News

“We have information that Moonshot distilled Fable for the development of K3”

twitter.com

471–480 of 742 posts

Re: “We have information that Moonshot distilled Fable for the development of K3”

#471

Earlier quoted context omitted.

distillation has been around for 12 years. it's not new in terms of ML techniques. https://arxiv.org/abs/1503.02531 although i doubt there has been a legal case over it yet in the context of the legality of stealing shit but IANAL.

Yes, I get that, but it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity. It's completey insane that we still don't know how Open Source would work, that the laws are vague and we're still technically waiting for the courts to decide on cases. The government should a) legislate and b) create test cases and run them through the courts so that we can have c…

> it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity.

you mean like the regulatory clarity surrounding stealing shit to make the LLMs in the first place?

> [There is] extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we understand how Fair Use applies to any one step in AI training, let alone all.

https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...

> The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.

if so, it would be nice if they approached the instances of stealing shit chronologically. but that's just my view.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#472
post #328

Does this matter? Distillation is not illegal by every definition of the word. There are millions of samples available on huggingface and models explicitely trained on output produced by fable. There has been no action taken against them. Another example is that it appears that the upper limit of what you can do is ultimately dependent on people working on the model, otherwise grok would be a LOT more competitive pre…

Of course it matters. Regardless of whether distillation is legal, there is a difference between training a model with and without distillation. For one thing, the distilled model wouldn't exist without the model it distilled. Also, companies that use distillation may be competitive but seem unlikely to surpass the companies that are training these models from scratch.

>but seem unlikely to surpass the companies that are training these models from scratch

Then why is it a problem?

Another serious question.

Trying to get my head around what the root of the objection is here. There must be some fear, but if that fear is not a fear of being surpassed in the market, then what is the fear?

Re: “We have information that Moonshot distilled Fable for the development of K3”

#473
post #265

Earlier quoted context omitted.

Perhaps even more importantly, the current frontier LLM models are self-admittedly the product of enormous quantities of copyright infringement and even less savory inputs, so calling them out for distilling the fruit of that tainted tree reads as highly hypocritical at best.

While I agree on a moral level, I think there is a distinction to be made. Training a SOTA model takes a huge amount of resources and expertise so the people doing the training are adding a lot of value along the way. I think this is much less true for distillation (which is kind of the whole point). ed: to clarify, I totally agree that a huge chunk of the value in LLMs is coming from the source material. My point wa…

Still, AFAIK Kimi's architecture (just like that of other LLMs from Chinese labs) is different from those of OpenAI and Anthropic's model in a nontrivial way. So the expertise is still there, and I guess resource use too (although Chinese labs tend to optimize this, thanks to the restrictions they have on GPU use).

EDIT: just wanted to add that resource optimization is usually where the contribution of Chinese labs is, so you shouldn't reaad the above parenthesis as a negative comment.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#475
Honestly, I don’t have any sympathy at all. Anthropic can complain all they want, but they seemed fine with pirating books. What Moonshot AI has done is to offer almost Fable 5 comparable performance at lower prices than Anthropic insane margins. This is what I call competition, which the Director seem to embrace.

This is what the Chinese always been good at. Take expensive innovation and streamline it to lower prices. But we are at a point where labs like Moonshot actually contributes a lot to the research field as well. They are pushing the innovation forward and squeezing the prices. Very well done.

Whats even weirder is the bizarre mechanisms Anthropic implemented to prevent distills which they had to sacrifice their customers for. They hid the internal CoT reasoning and returns summarizations instead. This made it difficult for users to trace things. They made Fable 5 silently switched over to Opus 4.8 if it detected blacklisted prompts (almost anything triggered this) to sabotage distills. And now, they are still complaining about distills? So their customers have gotten sacrificed over nothing.

Whats even weirder is the timeframe here, no way the Moonshot team managed to plan conduct a large scale distill, then pre-train, RL, fine-tune, benchmark, marketing and release to their platform since Fable 5 got whitelisted.

> they developed a sophisticated internal platform to conduct large scale distillation

I am very curious about this and would love to learn more on how they did this. Wish we had more details. I know the team behind DeepSeek have also done clever things to distill too. I am aware of these ”transfer stations” that acts as a proxy, but I don’t think they are helpful in this case.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#476
post #467

Earlier quoted context omitted.

The models were built using copyrighted works, so why can't models be built using other models?

Because model output is probably far closer to software or a licensed work which possibly has greater protections than it is to copyright. There is far less possibility of fair use, it might be protected by patents, license or reverse engineering laws. In any case the laws are being written now, but I doubt these will have worse protection than software does, which has far better protections than copyright

Why would this be the case. Why would software output from a model magically have greater protection than the software the model trained on.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#477

OpenAI and Anthropic should enter into distillation agreements with other US labs. Turn a threat into a profit center. Other US labs cannot directly distill from OpenAI/Anthropic as it’s a violation of the terms of service. It holds other US labs back. Leading them to build second tier models And in the end OpenAI/Anthropic may be unable to prevent distillation. Why fight it when there’s clear money to make here?

OpenAI/Anthropic are already charging for access to their closed models. They're getting paid.

They are also not interested in agreements. They want to keep as big a moat as possible because they love money. And you need two to tango.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#478
post #44
post #37

I can understand that the AI labs might care about other labs distilling their models as it can eat into their competitive advantage, but do consumers care at all? Aren't consumers benefiting from this practice by getting better cheaper models as a result?

They're probably going for the national security/domestic manufacturing angle. > Aren't consumers benefiting from this practice by getting better cheaper models as a result? Aren't consumers benefiting from cheap chinese batteries, EVs, and drones?

Heck yes!

I don't get why USA wants to keep losing money on sustaining failed uncompetitive zombie companies. The companies that decided to lose long term competence for short-term gains need to go bankrupt (you've had EV-1, but decided to drill, baby, drill). The greedy shareholders that rewarded destructive value extraction need to lose money, instead of getting a soft exit at taxpayers' expense.

If you want to give a subsidy, give it to something that will modernize and expand manufacturing, not to prolong death of companies whose entire R&D strategy is inventing new subscriptions for old car components.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#479
post #467

Earlier quoted context omitted.

The models were built using copyrighted works, so why can't models be built using other models?

Because model output is probably far closer to software or a licensed work which possibly has greater protections than it is to copyright. There is far less possibility of fair use, it might be protected by patents, license or reverse engineering laws. In any case the laws are being written now, but I doubt these will have worse protection than software does, which has far better protections than copyright

Let's assume model output can be claimed by copyright or some form IP. You can't really patent it, as the output isn't a novel idea or process, much like you don't patent a book or a movie. But for arguments sake, let's agree it is some kind of IP.

Who are you saying owns that IP? The people who trained the model? The people who ran the model? The people who wrote the prompt? The person who paid for all of that to happen?

If the model output is owned by the person prompting it and paying for the tokens, what's the problem here?

If the model output is owned by the trainer of the model, that's a big nasty can of worms.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#480

Earlier quoted context omitted.

If turning other peoples copyrighted work into a model is transformative enough to be protected then so is distilling that model into a different, better, model.

The models were built using copyrighted works, so why can't models be built using other models?

They do seem to be paying for it (as per the 1.5Bil lawsuit yesterday and them now purchasing books and licensing from media companies).

Whether we think they're paying enough is another question, but "I'm paying for content so can protect it" doesn't seem inconsistent.

We may decide that giving models away for free means they don't have to license content (judging by HN comments), but currently that doesn't seem to be the case as Meta is facing lawsuits for its open models.

(Obligatory stratechery piece: https://stratechery.com/2026/whos-afraid-of-chinese-models/ )

Post reply on HN