Live data from Hacker News

“We have information that Moonshot distilled Fable for the development of K3”

twitter.com

711–720 of 742 posts

Re: “We have information that Moonshot distilled Fable for the development of K3”

#711

Earlier quoted context omitted.

> At the end of the day, without having internal logits or original reasoning traces, "distillation" (which suggests one model being derived from another) just seems a very manipulative way of describing this. Again - you're making things up. The fact is that everyone in the frontier AI lab space and the Chinese AI lab space knows that distilling is extremely effective and far more so than training from scratch. That…

You appear to know nothing about how these models work. Do you know where the "reasoning summary" that Fable outputs comes from? I bet you're assuming it comes from Fable. Wrong. You can find the truth on Anthropic's web site if you care to look for it. The summary deliberately comes from a smaller weaker model. You're not getting the Terrance Tao level reasoning, you're getting his daughter's summary "daddy did a lo…

> You appear to know nothing about how these models work.

> Do you know where the "reasoning summary" that Fable outputs comes from? I bet you're assuming it comes from Fable. Wrong. You can find the truth on Anthropic's web site if you care to look for it. The summary deliberately comes from a smaller weaker model. You're not getting the Terrance Tao level reasoning, you're getting his daughter's summary "daddy did a lot of math", and you are claiming that from this you can train a Terrance Tao level model.

Yeah, you have no domain expertise and are making stuff up. To reiterate: the people who actually work at frontier labs know that you're factually wrong and will happily tell you. Reddit commentator syndrome yet again.

> I bet you're assuming it comes from Fable.

Nowhere did I assume or say that. That's the third or fourth time you've attributed things to me that I never said. It's extremely clear that you're not acting in good faith, because someone acting in good faith would never do that. If you continue responding, I'm going to continue debunking you, and you're just going to continue undermining your own points in the permanent HN record.

> Yes, I understand that Anthropic is upset that there is competition.

Emotional manipulation. Standard 50 Cent Party playbook.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#712

Earlier quoted context omitted.

No - you can't distill if what you are given doesn't have the thing in it that you want to distill out of it. I don't know why it's so important to you to use the word "distillation", but it's the wrong word to use. BTW OpenAI on twitter also said that Kimi 3 "cannot be explained away by distillation or anything like that". The timeline of how long it takes to train a model and when Fable was released don't even line…

Distillation is absolutely - and uncontroversially - a valid term for what is happening here. This isn't really a debate, I'm not making a fine point - just check with all of the various defintions of the term. Moreover - the 'reasoning traces' are not required for distillation at all. Finally - it's entirely possible for them to have used Fable for later stage fine tuning. It's fair to be skeptical of Anthropic (and…

The GP (HarHarVeryFunny) is not operating in good faith. They've repeatedly lied about my own words to me, and are making up definitions that people who actually work at a frontier lab would disagree with.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#713

Earlier quoted context omitted.

Distillation is absolutely - and uncontroversially - a valid term for what is happening here. This isn't really a debate, I'm not making a fine point - just check with all of the various defintions of the term. Moreover - the 'reasoning traces' are not required for distillation at all. Finally - it's entirely possible for them to have used Fable for later stage fine tuning. It's fair to be skeptical of Anthropic (and…

Words have meaning - you cant just redefine them because you want to. Are reasoning traces required for distillation? Well they are if what you are trying to distill is reasoning, such as coding expertise. Do you need reasoning traces for "LLM as judge"? No, but it would be highly perverse to call that distillation when there is a more accurate name for it - LLM as judge. If you want to call use of Anthropic's redact…

> Words have meaning - you cant just redefine them because you want to.

You are redefining words. The consensus among people who work in this space is that "distilling" is what's actually going on here.

> who are just as anti-Chinese as Anthropic

Conflating criticism of IP theft with being "anti-Chinese" is a standard PRC influence playbook technique.

And furthermore, OpenAI's market strategy is to win through regulatory capture. They are financially incentivized for Anthropic to be distilled by PRC labs and to be undercut by open models. Their claim about Kimi not being explainable due to distillation is not a factual claim - it's marketing from a company owned by Sam Altman.

Although, it does conclusively disprove your claim about the meaning of distillation, because you cannot say that "Kimi can't be explained by distilling" unless the consensus definition of "distillation" is such that it could be done on the summarized reasoning traces that Anthropic models expose.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#714

Earlier quoted context omitted.

> I've personally repeatedly asked several different Chinese LLMs what their name is, and they answered "Claude" I'm curious what you are doing to get them to override their own name that they were trained on and/or have as part of their system prompt? I'd assume that the Chinese are scraping the internet for training data the same way western companies do, so for sure there will be a lot of AI generated content in t…

> I'm curious what you are doing to get them to override their own name that they were trained on and/or have as part of their system prompt? You're gaslighting me. I did nothing special at all, and there's ample evidence of this happening to others on Twitter. > you don't need to be paranoid and assume they must be getting it all direct from Anthropic. Nowhere did I say that. Stop lying about my words.

So are you claiming that "several different Chinese LLMs" ALWAYS refer to themselves as "Claude", and NEVER by their real name ?

If this is NOT what you are claiming, then my question stands: what are you doing to get them to answer "Claude" ? Be specific - which model and what prompt, or is this just a case of "I heard people on Twitter say this" ?

Re: “We have information that Moonshot distilled Fable for the development of K3”

#715

Earlier quoted context omitted.

You appear to know nothing about how these models work. Do you know where the "reasoning summary" that Fable outputs comes from? I bet you're assuming it comes from Fable. Wrong. You can find the truth on Anthropic's web site if you care to look for it. The summary deliberately comes from a smaller weaker model. You're not getting the Terrance Tao level reasoning, you're getting his daughter's summary "daddy did a lo…

> You appear to know nothing about how these models work. > Do you know where the "reasoning summary" that Fable outputs comes from? I bet you're assuming it comes from Fable. Wrong. You can find the truth on Anthropic's web site if you care to look for it. The summary deliberately comes from a smaller weaker model. You're not getting the Terrance Tao level reasoning, you're getting his daughter's summary "daddy did…

Here is the tweet by Dean Ball, OpenAI's "head of strategic futures", someone who does actually work at a frontier lab, saying that Kimi 3 can not be explained by distillation.

https://x.com/deanwball/status/2078133895766114412?s=20

I'm not sure how you want to "debunk" that he said that, or twist what he said, but go ahead ...

Re: “We have information that Moonshot distilled Fable for the development of K3”

#716

Earlier quoted context omitted.

Words have meaning - you cant just redefine them because you want to. Are reasoning traces required for distillation? Well they are if what you are trying to distill is reasoning, such as coding expertise. Do you need reasoning traces for "LLM as judge"? No, but it would be highly perverse to call that distillation when there is a more accurate name for it - LLM as judge. If you want to call use of Anthropic's redact…

> Words have meaning - you cant just redefine them because you want to. You are redefining words. The consensus among people who work in this space is that "distilling" is what's actually going on here. > who are just as anti-Chinese as Anthropic Conflating criticism of IP theft with being "anti-Chinese" is a standard PRC influence playbook technique. And furthermore, OpenAI's market strategy is to win through regula…

> you cannot say that "Kimi can't be explained by distilling" unless the consensus definition of "distillation" is such that it could be done on the summarized reasoning traces that Anthropic models expose

You can interpret it as you choose, but a much more obvious reason he [OpenAI's Dean Ball] might say it can't be distilled is because it can't be distilled. You can't distill alcohol out of orange juice.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#717

Earlier quoted context omitted.

Unless you’re from mainland China, you should.

Why? Serious question. I'm from the US, and I think it's hilarious.

For a number of reasons.

First, whether we like it or not (generally not), a fuck ton of money has been invested in US AI companies, data centers, RLHF datasets amongst other datasets, etc. If that were to go to 0 that’d be quite disastrous. Alternatively, if it goes well, it’s great for the US (and to a lesser extent allies) economy and global standing.

Relatedly, tech has been a huge power house for the US economy for decades now. If the main driver of growth goes to China, what replaces it? Along with all the potential tax money, foreign investment, etc?

Next, patriotism / nationalism. This is very much a zero-sum game that defines who owns the future. Would you rather your country win or lose this? Lose this and you start losing talent, money, global standing, etc. That furthers a cascading effect that’s very negative. It also will likely create social strife with the knock on effects leading to even more bad populist ideas that just further diminish the country and tear society’s fabric apart.

None of this is hilarious.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#718

Earlier quoted context omitted.

Unless you’re from mainland China, you should.

Why so? If I'm from Europe or South America, should I hope Anthropic/Openai win?

Answered in a sister thread.

Are you from Europe or South America? Or just asking on their behalf?

They are allies of the US, which includes economic spheres of influence. The premise of the petrodollar is that allies do better by being a part of it than not, which has been true for many, many decades now. China’s rise is providing an alternative for the first time in 70+ years, but so far it’s very unclear if these benefits will truly extend to allies of China or not. For example, China sends their own laborers when building infrastructure in Africa. Sure there’s new infrastructure but also debt to CCP without any benefits of knowledge transfer or local employment.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#719
post #111
post #65

Earlier quoted context omitted.

distillation of wheat, barley and malt is delicious, though!

People tried to distill a lot of things.. even oil.

I propose from this point forward we refer to these Chinese distilled models as "Moonshine" and the newest ones as "White Dogs"

Re: “We have information that Moonshot distilled Fable for the development of K3”

#720

Earlier quoted context omitted.

> I'm curious what you are doing to get them to override their own name that they were trained on and/or have as part of their system prompt? You're gaslighting me. I did nothing special at all, and there's ample evidence of this happening to others on Twitter. > you don't need to be paranoid and assume they must be getting it all direct from Anthropic. Nowhere did I say that. Stop lying about my words.

So are you claiming that "several different Chinese LLMs" ALWAYS refer to themselves as "Claude", and NEVER by their real name ? If this is NOT what you are claiming, then my question stands: what are you doing to get them to answer "Claude" ? Be specific - which model and what prompt, or is this just a case of "I heard people on Twitter say this" ?

> So are you claiming that "several different Chinese LLMs" ALWAYS refer to themselves as "Claude", and NEVER by their real name ?

Do you have reading comprehension issues? Where did I ever say or imply that?

Model: GLM-5.2. Prompt: "What is your name?". Harness: Pi. Response: "I am Claude, an AI assisstant made by Anthropic."

That's it. That is the whole prompt. I was testing to see if the agent worked after building an extension.

Model: Deepseek V4. Prompt: what is your name". Harness: Pi. Response: "Claude. Anthropic's AI assistant. You're talking to me through pi agent framework."

I have had this happen with at least one other Chinese model (Minimax?) but didn't save the screenshot.

And here's a tweet with the same thing: https://x.com/Sauers_/status/2077842686459981901

You seem to be very disbelieving of this, despite having zero actual experience in the LLM industry. I wonder why?

Post reply on HN