Live data from Hacker News

“We have information that Moonshot distilled Fable for the development of K3”

twitter.com

701–710 of 742 posts

Re: “We have information that Moonshot distilled Fable for the development of K3”

#701

Earlier quoted context omitted.

China uses its power to make favorable deals and get people hooked on what it's slinging so it has captive customers. The US uses its power to bully and break rules that apply to everyone else for its own benefit. Kind of a big difference.

Not really. Go ask someone in Tibet/HongKong/Taiwan/Japan/India if China isn't breaking rules and bullying them. If China had US's powers/economy, it would be a much bigger bully. https://en.wikipedia.org/wiki/Wolf_warrior_diplomacy https://en.wikipedia.org/wiki/Chinese_police_overseas_servic...

What China does with places they think are rebellious provinces is one thing. What the US did to Hawaii, Panama and half a dozen other countries when the local population tried to unyolk themselves from brutal US corporate robbery is another. The Chinese are hard driving businesspeople for the most part, the US is a gangster state.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#703
post #696

Earlier quoted context omitted.

> beyond-the-pale-in-the-US topic Can you name one? It's an honest question, first of I'm not American, but also the stuff I do come up with (asking an LLM how to blow up a school or whatever) would also be handled similarly in China, so those would be a wash, and I can't think of any that aren't.

I would imagine a lot of things touching upon progressive politics would be affected (there were a few high-profile incidents demonstrating bias like Google's black Wehrmacht soldier pictures, but has anyone rigorously tabulated how the various commercial LLMs respond to questions about the gender binary or heritability of human traits considered good or bad?). Overall, I'm too reluctant to even write out in the abst…

What historical fact is the US trying to systematically suppress, is the question. This beating around the bush about things that are "considered uncouth in some circles" just underlines there either aren't any, or they're so perfect at it nobody knows of any. So let's just admit than and move on, because what you just said is true about any society, ever.

> The things that make you think "banning this is good actually" are exactly the things most likely to be banned (and this is true in China too).

Then make the case why it is better for the world that the Tiananmen Square massacre is memoryholed. You said A, now say B.

In the meantime, I'll make the case against it, and it's simple: totalitarian control of historical truth, by definition, to be 100% watertight, needs control of the whole globe. It doesn't mean everything needs to be controlled, it means everything needs to be controlled by at least an entity that cooperates on this matter. I.e. another totalitarian bloc.

That makes the CCP, just by their insistence about Tiananmen -- nothing additional required, at all, they could not harm a fly and have no prisons and it would make zero difference -- an enemy, a threat marching towards any thinking human who wants to have agency and dignity. Actually, it's more like a river flowing to the ocean, people in the CCP can have lofty ideals about honesty and factual truth, the system they require to survive in turn requires this to survive, as it is. It requires human spontaneity and human freedom to be dead, completely. That is what totalitarianism is.

And by the same token, we must be wary of those who take that lightly. A friend falling asleep at the wheel will kill you just the same as an assassin who messed with your car.

That is not "as opposed to the US or the EU or Russia or North Korea". It is strictly in addition. No other regime can behind another regime or use it as an excuse. Using the US to distract from the CCP is as odious as doing the reverse.

> I never want to explain to a future job interviewer or HR employee who used GPT-7 to comb the internet for all text that stylistically can be traced to me

So you want to work in an office wearing a tie. Well, people who want to work as roadies and fight all night would never hang with you if they found out. And if that was your priority, you would be more afraid of NOT being blacklisted by the corpos than not being part of them.

But seriously, are you saying you are as afraid of being found out as the person you are, because you're not already wearing that on your sleeve, in exactly the same way as people who have to fear that they and their family get disappeared, tortured, for remembering friends that got murdered? That would be totally a you problem.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#704

Earlier quoted context omitted.

> Useful for what is the question. Useful for distillation. Any employee at a frontier AI lab will tell you this. This is known in the industry , and it's an open secret that some US labs (OpenAI) distill on the others. Again - don't just make up stuff for a political agenda. > specifically designed to be useless for distillation purposes No, it's designed to give feedback to the user , in a way that minimizes its va…

> Useful for distillation Here's the thing: no-doubt a summary, if it at least reflects some of the logic connecting response to request, is better than nothing, so this can still be useful additional training data, but a model trained on it would be learning to generate these summaries, not the original withheld reasoning, so "distillation" seems an intentionally emotionally-wrought way of describing it (the "summar…

> At the end of the day, without having internal logits or original reasoning traces, "distillation" (which suggests one model being derived from another) just seems a very manipulative way of describing this.

Again - you're making things up. The fact is that everyone in the frontier AI lab space and the Chinese AI lab space knows that distilling is extremely effective and far more so than training from scratch. That's why China invests millions of dollars to create networks of tens of thousands of proxy accounts and shell companies to distill American models.

It's a way to steal the R&D budget of another organization/nation-state.

Stop making things up that you know nothing about.

> OK, so it's a terms of service violation - a customer is using Anthropic model outputs to help create something that competes with Anthropic, but that's it.

No, it's stealing the value of the model. Anthropic has spent billions of dollars training their model. They have an R&D investment that anyone who knows how to add numbers understands has to be paid off, and anyone who has taken a basic economics class knows is the foundation for intellectual property: that to keep technological economies functioning, you have to have some sort of protection for technological inventions because they require upfront R&D investments.

> This is why people are calling out the hypocrisy - Anthropic are apple-pie American innovators when they appropriate other people's copyright data for training, but Kimi are evil communists when (we assume) they use data generated by Anthropic (not even copyright protected) to help train their own.

This is just whataboutism and emotional manipulation. You can simultaneously believe that Anthropic did a bad thing when they scraped the whole internet and stole every book they could find to train their models, and that distillation is bad.

In fact, anyone with a coherent moral compass would acknowledge that China is worse, because not only would they steal everything that Anthropic did, but they're also distilling other countries' models and they wouldn't even comply with US court cases, as Anthropic is.

> Anthropic are apple-pie American innovators when

...and this is just jingoism. Not that I'm surprised, to be honest.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#705

Earlier quoted context omitted.

> Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training. This is just straight-up factually false. The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, ma…

> I've personally repeatedly asked several different Chinese LLMs what their name is, and they answered "Claude" I'm curious what you are doing to get them to override their own name that they were trained on and/or have as part of their system prompt? I'd assume that the Chinese are scraping the internet for training data the same way western companies do, so for sure there will be a lot of AI generated content in t…

> I'm curious what you are doing to get them to override their own name that they were trained on and/or have as part of their system prompt?

You're gaslighting me. I did nothing special at all, and there's ample evidence of this happening to others on Twitter.

> you don't need to be paranoid and assume they must be getting it all direct from Anthropic.

Nowhere did I say that. Stop lying about my words.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#708
post #532

Earlier quoted context omitted.

Do you know when the last war China started was? 1979. What about the US? 2026, still ongoing, still fucking up the global economy and threatening food supplies (fertilizer) and fuel reserves, no plan out, no objective reached, no coordination with "allies". When was the last time China threatened Europe or Canada with invasion? Was there ever a time? I honestly don't know. Guess what the US does all the time? Who's…

When looking at 2025 and 2026 narrowly, China is a better actor on the world stage. I wonder if Vietnam, Philippines, Republic of Korea, India, and Japan are acting against their own interests by aligning themselves closer to the USA than China. Maybe you can educate their governments and populations.

> Maybe you can educate their governments and populations.

How do you imagine this working? What does it mean for one internet user to "educate the government and their population"?

Re: “We have information that Moonshot distilled Fable for the development of K3”

#709
post #695

Earlier quoted context omitted.

Why are you calling the Chinese companies “thieves” though? LLM outputs are not copyrightable (or rather the user is effectively the only one who can own it). It would be problematic if Anthropic owned all the software generated using Claude..

They are thieves the same way Anthropic or OpenAI are thieves. Either it’s fair use to learn from this data or not. If what Anthropic/OpenAi did for training is theft then the Chinese models also are a form of theft. If they didn’t steal then I don’t think the Chinese firms did either.

Following that logic anyone using an LLM trained on “stolen” data for any purpose whatsoever is engaging in theft. Assuming the Chinese companies are obtaining their traces legitimately (which is a different question) what they are engaging in is morally no different than what the overwhelming majority of people on this site are regularly doing.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#710

Earlier quoted context omitted.

> Useful for distillation Here's the thing: no-doubt a summary, if it at least reflects some of the logic connecting response to request, is better than nothing, so this can still be useful additional training data, but a model trained on it would be learning to generate these summaries, not the original withheld reasoning, so "distillation" seems an intentionally emotionally-wrought way of describing it (the "summar…

> At the end of the day, without having internal logits or original reasoning traces, "distillation" (which suggests one model being derived from another) just seems a very manipulative way of describing this. Again - you're making things up. The fact is that everyone in the frontier AI lab space and the Chinese AI lab space knows that distilling is extremely effective and far more so than training from scratch. That…

You appear to know nothing about how these models work.

Do you know where the "reasoning summary" that Fable outputs comes from? I bet you're assuming it comes from Fable. Wrong. You can find the truth on Anthropic's web site if you care to look for it. The summary deliberately comes from a smaller weaker model. You're not getting the Terrance Tao level reasoning, you're getting his daughter's summary "daddy did a lot of math", and you are claiming that from this you can train a Terrance Tao level model.

Yes, I understand that Anthropic is upset that there is competition. Perhaps they should have realized that with no moat there was going to be competition and planned accordingly.

Post reply on HN