Earlier quoted context omitted.
"Why would other countries, that don't share the same anxiety about China as the US, would be troubled with the this?" It's the other way around. There is a high likelihood that many countries of the "west" (the "global north"?) will outlaw, restrict, or otherwise control LLMs and the tools that enable them. The US, however, is blessed with the first amendment which makes it extremely difficult to restrain speech in…
It wasn't difficult for the US to restrict TikTok (or BYD, Huawei, DJI...)
The Kimi K3 Moment
561–570 of 644 posts
Re: The Kimi K3 Moment
#562Earlier quoted context omitted.
How is distillation an "attack" but gigascraping the Internet to the point of crashing servers and everyone needs Cloudflare and Anubis now not an "attack"? I'm not aiming for a what about kickflip here: I'm saying we need to either agree on some rules or stop crying foul. Maybe the coherent legal theory is that neural networks and intellectual property don't interact. That would be weird but it would be consistent,…
I think you've basically got the legal theory. Training a neural network isn't prohibited by copyright law so if you can legally get your hands on something (e.g. by sending a GET request to someone with rights to serve the contents of their web page, or by buying a book) without signing a contract to not train on it, you can train on it. But the American AI companies only let you query their models if you first sign…
Re: The Kimi K3 Moment
#563Earlier quoted context omitted.
Contract law is never going to prevent this.
Why not? Seems like a perfectly normal contact term to me. Or do you just mean that US courts don't have enough teeth to prevent Chinese companies from violating contracts? On that I agree.
Re: The Kimi K3 Moment
#564Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…
> Distillation “attacks” are not attacks. Say it louder for the people in the back. All these complaints about "distillation" from frontier labs are bordering on felony contempt of business model at this point. It's great for us. Maybe it's bad for them but nobody other than shareholders really cares. The optimal outcome for humanity is for oligarchs to spend trillions training a godlike AI, only for the precious wei…
Re: The Kimi K3 Moment
#565Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…
> Distillation “attacks” are not attacks. If "distillation attacks" happen, we have to conclude there is some value add in what model labs do. Regardless of how we feel about using existing human knowledge in the way they currently do, it's simply impractical to infer that everything that happens downstream of LLMs can not be an attack on some IP because of it. So both things can be true: a) People infringe on Anthro…
Re: The Kimi K3 Moment
#566Earlier quoted context omitted.
> This behavior is exactly what you'd expect from a model distilled from Claude. This is not at all what I would expect because it's trivial to change the training data to replace Claude with Kimi. In fact I'd argue it's almost certainly not saying that due to distillation.
>This is not at all what I would expect because it's trivial to change the training data to replace Claude with Kimi. Wait what? The reason you wouldn't expect it is because if it was distilled, it would be easy to get rid of self identification? Is that any less true of a non distilled model? I suppose there's lots of ways to interpret it, but the idea that self-identifying as Claude is affirmative evidence that it'…
I wasn't arguing that. I was arguing that even if it was distilled from Claude, the distillation isn't why it identifies as Claude. Therefore identifying as Claude isn't evidence of distillation.
Claude has been caught identifing itself as Deepseek:
https://news.ycombinator.com/item?id=47145081
I don't take that to mean it's necessarily been distilled on Deepseek.
Gemini was supposedly caught identifying as Claude:
https://www.reddit.com/r/ChatGPT/comments/1gslm0t/gemini_mod...
I don't take that to mean it was distilled on Claude.
Claude was caught identifying as ChatGPT:
https://www.reddit.com/r/OpenAI/comments/1e34tkr/why_is_clau...
I don't take that to necessarily mean it was distilled on ChatGPT.
Re: The Kimi K3 Moment
#567Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…
At this point it may not even be happening intentionally given the quantity of LLM-generated content that is appearing online and is likely being re-ingested by models.
Re: The Kimi K3 Moment
#568Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…
I strongly agree with the premise that distillation is not an “attack”. But that said: K3 is not a distilled version of Fable or Sol. Fable has been barely available and Sol was just released! Moreover, K3 is superior to both models in some domains, according to user scoring on the Arena. API distillation can’t give you these results anyway. All it is useful for is bootstrapping RL in new domains to get past the “col…
Don’t you think there’s maybe a teeny, tiny chance that their approach is a little more sophisticated than just buying retail subscriptions? Maybe the trade secrets are being exfiltrated directly from the American labs?
Re: The Kimi K3 Moment
#569Earlier quoted context omitted.
The china talk page points to an article about crypto currency stolen credit cards. If I believed every random X account i would have to believe too many false things. Not really convinced my guy.
The companies running these proxy stations are already: 1) Using botnets to mass-create thousands of accounts 2) Blatantly violating ToS by splitting and reselling accounts 3) Creating thousands of accounts using fraudulent identities 4) Bypassing KYC by recruiting real people in low-income countries for biometric face-matching checks for a few dollars 5) Using AI deepfakes to fake passports / verification credential…
Re: The Kimi K3 Moment
#570Well, there is the small issue of privacy policy: Kimi will train their models on your interactions if you use their subscriptions, and only with direct API usage (billed at API prices) they say they won't. Whether you trust that is another matter. Those things do make a difference to some of us, even though nothing is black and white. In my case, I'll probably want to wait until other providers appear through OpenRo…
> Kimi will train their models on your interactions I find these kinds of concerns increasingly silly: most of the input to these models will be ... previous output from the very same models, alongside the occasional half-assed human command to fix something and "make zero mistakes". Who cares if they train on that? Let them, if it makes their future models better! 99% of users are not working on any special IP to wo…
I haven't measured the percentage of users, so I wouldn't venture an opinion with a number, but these concerns do exist.