Live data from Hacker News

“We have information that Moonshot distilled Fable for the development of K3”

twitter.com

681–690 of 742 posts

Re: “We have information that Moonshot distilled Fable for the development of K3”

#681

Earlier quoted context omitted.

> Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training. This is just straight-up factually false. The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, ma…

Well there were observed cases of Claude calling itself Deepseek or Qwen. So pot calling the kettle black? To be fair I find it hard to take this too seriously, shouldn’t it be trivial to just replace “Claude” with any other string in your “distillation” dataset?

> Well there were observed cases of Claude calling itself Deepseek or Qwen. So pot calling the kettle black?

There are open-source Deepseek and Qwen models - "distilling" doesn't involve breaking terms of service or hitting an API because you can literally run local inference or even just inspect the weights directly, and that's intended because they're open source.

It's categorically different for a nation-state to build massive illicit networks of fraudulent identities to do distillation over tens of thousands of accounts to intentionally bypass providers' terms of service, intention for their models, and business model that very explicitly proprietary and not open source.

https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...

If Claude did distill on proprietary PRC LLMs - then fine, shame on them - I condemn that and I expect others to do the same. But there are no open-source Claude models. The only way for PRC models to have those responses is if they distilled Anthropic's models from their APIs.

> To be fair I find it hard to take this too seriously, shouldn’t it be trivial to just replace “Claude” with any other string in your “distillation” dataset?

...and what would happen when it read all of the books and articles about Anthropic and replaced "replaced Claude Opus" with "replaced Qwen Opus"? Did you give any thought to this at all before saying it?

Re: “We have information that Moonshot distilled Fable for the development of K3”

#682

Earlier quoted context omitted.

> Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training. This is just straight-up factually false. The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, ma…

> The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable. Useful for what is the question. Nobody is debating whether the outputs of LLMs are valuable. Given that Anthropic have redacted their true reasoning, and replaced it with a "summary", specifically designed to be useless for di…

> Useful for what is the question.

Useful for distillation. Any employee at a frontier AI lab will tell you this. This is known in the industry, and it's an open secret that some US labs (OpenAI) distill on the others. Again - don't just make up stuff for a political agenda.

> specifically designed to be useless for distillation purposes

No, it's designed to give feedback to the user, in a way that minimizes its value for distilling. It's still valuable, and so there's a good chance that they'll remove it entirely as a result.

> it would certainly be highly ironic if this summary was in fact "even more valuable" for that purpose as you are claiming!

I did not claim that. Read my comment again:

> The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable.

Because apparently I have to spell it out:

The output of a reasoning model is valuable, even if it didn't have the reasoning summary. Anthropic's models have a reasoning summary. The reasoning summary makes the output more valuable than if it didn't have a reasoning summary. It does not make it more valuable than having the full reasoning.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#683

“However, large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable.” What’s actually happening behind the scenes is that certain inference providers will classify a prompt and it’s re-routed transparently to Anthropic and that’s used for distillation training, only distilling the complicated traces they need, originating from real user…

> What’s actually happening behind the scenes is that certain inference providers will classify a prompt and it’s re-routed transparently to Anthropic and that’s used for distillation training

Uh, no. There are Chinese networks of tens thousands of fake identities specifically to get access to Anthropic models directly.

https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...

Don't make up stuff and/or lie to suit a political agenda. It's extremely dishonest.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#684
post #622

[flagged]

Can you please not fulminate or post flamebait on HN? This is in the site guidelines: https://news.ycombinator.com/newsguidelines.html . You may not owe AmericanAIBros better but you owe this community better if you're participating in it.

I’m familiar with the guidelines and I stand by what I said, including how I said it. This has been a regular talking point from AI companies since DeepSeek first hit the scene, and I feel a glib response to what is very clearly an insincere and nakedly hypocritical talking point is warranted after years of this slop.

Hypocrisy doesn’t warrant professionalism, it warrants corrective action; in text, the best I can offer is a tone and tenor that matches the original argument. Considering these dolts have now made the claim that open weights somehow equates to AI communism, this sort of response is even more necessary than before to reflect the complete absence of decorum from the people making these grievances in the first place.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#685
post #532

Earlier quoted context omitted.

Do you know when the last war China started was? 1979. What about the US? 2026, still ongoing, still fucking up the global economy and threatening food supplies (fertilizer) and fuel reserves, no plan out, no objective reached, no coordination with "allies". When was the last time China threatened Europe or Canada with invasion? Was there ever a time? I honestly don't know. Guess what the US does all the time? Who's…

When looking at 2025 and 2026 narrowly, China is a better actor on the world stage. I wonder if Vietnam, Philippines, Republic of Korea, India, and Japan are acting against their own interests by aligning themselves closer to the USA than China. Maybe you can educate their governments and populations.

[flagged]

Re: “We have information that Moonshot distilled Fable for the development of K3”

#686

how is it possible to distill fable only a month after its release? maybe they are confusing opus with fable.

Create a couple thousand Claude max accounts and split the work amongst them perhaps.

...which they already had:

https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...

The line "how could they do it in such a short time" is absolutely idiotic. The infrastructure was already there, it's massively parallel, and it's not like the only thing being distilled on is Fable.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#687

Earlier quoted context omitted.

> beyond-the-pale-in-the-US topic Can you name one? It's an honest question, first of I'm not American, but also the stuff I do come up with (asking an LLM how to blow up a school or whatever) would also be handled similarly in China, so those would be a wash, and I can't think of any that aren't.

From an outsiders point of view: - The Trail of Tears - The Tuskegee syphilis study - Use of Agent Orange in the Vietnam War - Open Air biological warfare testts in civilians eg. in 1950 San Francisco - The only use of nuclear weapons against civilians? - Coca Cola and "american culture" - Neoliberalist economy - Spreading blame for their sins to other "white" nations plus one: The text input method to HN comments :(

all of these things have books about them published in the US

you are comparing a wooden stick to a fighter yet, try again

Re: “We have information that Moonshot distilled Fable for the development of K3”

#688

Earlier quoted context omitted.

You're right, I didn't include stuff at the beginning, for example, the theft of IP in textile manufacturing in the late 1700s. The US government didn't recognize copyrights on foreign literature which let US publishers reprint things such as Gilbert, Sullivan's operettas and Dickens novels Then there is Alexander Hamilton's advocacy for importing foreign technicians that bring back IP and reproduce it here in the st…

Which is to say they didn't have much of an idea at all, because it really didn't exist in much the same way. In fact, this is still a inaccurate characterization for exactly that reason. On the basis of copyright, take for instance the idea of exclusive rights to print a work. This actually wasn't implemented as a method of protection for the author, but a political reaction to the printing press being "misused" in…

I guess I should thank you for adding yet another load of deeper reading into American history. :)

Re: “We have information that Moonshot distilled Fable for the development of K3”

#689

Earlier quoted context omitted.

> The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable. Useful for what is the question. Nobody is debating whether the outputs of LLMs are valuable. Given that Anthropic have redacted their true reasoning, and replaced it with a "summary", specifically designed to be useless for di…

> Useful for what is the question. Useful for distillation. Any employee at a frontier AI lab will tell you this. This is known in the industry , and it's an open secret that some US labs (OpenAI) distill on the others. Again - don't just make up stuff for a political agenda. > specifically designed to be useless for distillation purposes No, it's designed to give feedback to the user , in a way that minimizes its va…

> Useful for distillation

Here's the thing: no-doubt a summary, if it at least reflects some of the logic connecting response to request, is better than nothing, so this can still be useful additional training data, but a model trained on it would be learning to generate these summaries, not the original withheld reasoning, so "distillation" seems an intentionally emotionally-wrought way of describing it (the "summary" is generated by a different smaller model - not the one the rest of the response came from).

It does bring up an interesting point though - RL training in general results in "cargo-cult" reasoning - you train a model to follow the steps (mistakes and all - Karpathy) that got to a result, without understanding why they worked. If this training on summaries works just as well as training on detailed reasoning, then it just highlights how having a few breadcrumbs to follow/regurgitate is all that it takes.

At the end of the day, without having internal logits or original reasoning traces, "distillation" (which suggests one model being derived from another) just seems a very manipulative way of describing this. OK, so it's a terms of service violation - a customer is using Anthropic model outputs to help create something that competes with Anthropic, but that's it. They are not copying Anthropic - they are, one assumes, using other models to generate cheap training data that they would otherwise have to pay people to generate.

This is why people are calling out the hypocrisy - Anthropic are apple-pie American innovators when they appropriate other people's copyright data for training, but Kimi are evil communists when (we assume) they use data generated by Anthropic (not even copyright protected) to help train their own.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#690

Earlier quoted context omitted.

It is incredibly important to whether the US can maintain its AI lead. If foreign competition is closing the gap only by distillation, then the frontier labs can focus on preventing distillation and maintain their lead that way. US dominance is also important for approaches to safety, especially political approaches. If the frontier models are all US-based, safety might be tackled via internal US policy. If other cou…

Do Americans even believe that US policy is likely to steer development in a way that’s safe and beneficial for humanity? The rest of the world certainly doesn’t. The US currently seems to primarily use their superpower status to be the world’s number one shit disturber and geopolitical antagonist. I don’t think China’s necessarily any better, but I’d rather have the most powerful models be open rather than under the…

I didn't mean to imply that the US is more likely than elsewhere to responsibly steer AI via policy. But I do think it is easier if it can be done internally as opposed to via international dealmaking.
Post reply on HN