Earlier quoted context omitted.
We're still in the early days of the AI industry timeline(relative to traditional industries). Not everything has yet been litigated. Taxes on AI subscriptions or AI capable hardware, to financially compensate IP holders for (potential) IP theft, could very well arrive in the near future, once the industry is mature. If this shocks you and sounds preposterous, I'll remind you that in several EU countries, we still pa…
I think we are going to direction where AI corps will have stronger lobby compared to IP holders.
The Kimi K3 Moment
321–330 of 644 posts
Re: The Kimi K3 Moment
#322Earlier quoted context omitted.
While it sounds like a lot, do you suppose 3.4 million sessions come even close to being sufficient to train a frontier model? Assuming each session was 10,000 words each, that's 34 billion words; lets call it 50 billion tokens (0.05 trillion) unfairly pilfered from Claude. That left Moonshot needing to scrounge for the other 14.950 trillion training tokens required for a baseline frontier model.
3.4 million is the number of sessions Anthropic detected. The actual number of Claude sessions trained on is likely >100 million. There are tens of thousands of accounts funneling Claude sessions into Chinese labs https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens... They are used for post-training, i.e. calibrating the model to understand and use tools/command line more effectively.
That's an increase of only a single order of magnitude, increasing my estimate of exfiltrated tokens from 0.05 to 0.15 trillion - a far cry from the 15 trillion required.
> They are used for post-training
Possibly - it may be too much data for post-training, unless further curation was done. However, this is not distillation; you know it, I know it, Dario knows it, but "Distillation Attack" is a short, memorable, sciencey-sounding, political sound-bite with enough malevolence to be deployed on the floors of congress, or by the usual fear-mongering newstainment talking heads.
Re: The Kimi K3 Moment
#323Earlier quoted context omitted.
Because the transformer architecture that enabled modern LLMs wasn't invented until 2017[1]? 1: That's the "T" in GPT fyi, even though Google is the author of the research paper that changed everything
Right. So we had enough information to train LLMs but not the technology to build it. So the initial models arent just distilled from information. We’ve always had the information.
Re: The Kimi K3 Moment
#324Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…
Re: The Kimi K3 Moment
#325Earlier quoted context omitted.
> People infringe on Anthropics IP Unless someone literally stole the weights somehow (which is not out of the question, I doubt either oAI/Anthropic have the capabilities to prevent a state-level actor getting those weights), distillation from generations is not infringement on anyone's IP nor is it stealing nor is it an attack. It can't be. As long as you pay for tokens you get to do whatever you want with them. So…
Its definitely an attack. Thats established from anthropics perspective. No one has a right to use Anthropic’s services in ways that directly violate the ToS and user agreements.
Re: The Kimi K3 Moment
#326Earlier quoted context omitted.
We really need to stop using $/M tokens as the pricing benchmark. I've found that the number of tokens used tends to be a bigger factor than the listed per token price. The cost per task vs. intelligence curve is really what you care about, and in my estimation Chinese models are just not there. They are focused on benchmaxing and getting the highest raw score they can, rather than efficiency.
Yes, this is already accounted for in many benchmarks, but without deep context of the problem type, the top line pricing is the best starting point. In my own experience, Fable is more token efficient than opus 4.8 with a higher likelihood of completing tasks correctly or at least with minimal corrective work. Opus regularly struggled to gather the correct context and reason effectively about what it had gathered. G…
Re: The Kimi K3 Moment
#327Earlier quoted context omitted.
Does it matter? As an end user I really only care about 1) how much I can do in a week, and 2) how long each task takes. Subsidies would affect 1, but not 2. But if some VC wants to subsidize my Claude or Codex or whatever, awesome.
The more important question than subsidy is what is the tokenomics of running the model. If it's inefficient to run on an nvl72 cluster (or whatever the heck has enough vram to run a 3T parameter model), and k3 isn't very token efficient, then it might not be that compelling of an open weights model.
Re: The Kimi K3 Moment
#328Earlier quoted context omitted.
Almost all markets depend on some form of regulation whether its as simple as "leave everyone alone but no stealing" or "every participant has to source every object through mountains of red tape." Thus far the US has not really chosen to go the Chinese rare-earth method yet. The problem with distillation attacks is the end result is everyone who is not doing them is going to deal with some kind of regulation whether…
Given these models could not have been trained in the first place if they had to license every line of random fan fiction on the internet, I think distillation also being fair game is a tradeoff everyone should be willing to take (unless they want to decelerate, but that's a different conversation).
Re: The Kimi K3 Moment
#329Re: The Kimi K3 Moment
#330Earlier quoted context omitted.
Absolutely do not pay for the kimi plans thinking they will be cheaper. If you sign up with a Chinese phone number, you can get the same plan for 200 yuan instead of 200 usd, it also only accepts Chinese payment methods iirc. So the plans are really made for Chinese userbase.
Wow! Does it accept a foreign alipay/wechat pay account?
... but I borrowed a friend's +86 phone number, which you'll need to even see that price. or maybe a 回国 VPN will work.