Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…
The Kimi K3 Moment
491–500 of 644 posts
Re: The Kimi K3 Moment
#492I tried Kimi K3 on a task I've done with every other model I use regularly ( https://swelljoe.com/post/i-let-every-agent-implement-its-ow... ) and found it chewed a lot longer on the problem and ate up almost the entirety of a 5 hour usage limit on their $19 plan. I only have the $20 plan from OpenAI and the same task, with a lot of the same implementation details as Kimi Code, only took a few minutes and consumed al…
We really need to stop using $/M tokens as the pricing benchmark. I've found that the number of tokens used tends to be a bigger factor than the listed per token price. The cost per task vs. intelligence curve is really what you care about, and in my estimation Chinese models are just not there. They are focused on benchmaxing and getting the highest raw score they can, rather than efficiency.
Unless the cost per token is prohobitedly high, people can often try the model out themselves and make a subjective judgement of how effective and efficient is it at solving tasks they usually deal with, using their setup.
Re: The Kimi K3 Moment
#493Earlier quoted context omitted.
Us models didnt pay for licenses too
That is incorrect. Anthropic paid $1.5 billion in compensation to copyright holders for use of their content in training data. OpenAI pays hundreds of millions per year across 150+ licensing deals for access to copyrighted data. Meta and Alphabet have similar arrangements. Under the settlement, Anthropic was forced to delete the pirated data they were training on. Chinese labs can still train on pirated data. I doubt…
Re: The Kimi K3 Moment
#494This was always where this was heading, but we got here much faster than expected. Once western governments declare it to be a "national security" risk for citizens to have access to open-weight frontier models, and once they classify using these models as acts of terrorism, what will that world be like? Will using Kimi K3 come to be like how napster was in the olden days? Everybody knew it was technically illegal, b…
even 8x rtx pro 6000 is only 768GB of VRAM. IDK how anyone is going to run k3
Re: The Kimi K3 Moment
#495Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…
I wonder, how does distillation deal with unprobed spaces in the knowledge landscape? Is a distilled model worse in some niche area that was not probed? Presumably, this is why frontier labs dont distill their own models internally to release them to the public as a servicable frontier model.
Re: The Kimi K3 Moment
#496Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…
Re: The Kimi K3 Moment
#497Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…
The desire to accuse China of just copying is like 20 years out of date. It’s been wrong since some people on HN were in diapers. People are going to be gobsmacked when, in our lifetime, China becomes a world power comparable to the U.S. Probably still poorer per capita, but at Spain/Italy levels, not third world country levels. And they’ll be shocked at the implications of that on the world economy, migration patter…
Re: The Kimi K3 Moment
#498Earlier quoted context omitted.
Does anything about that strike you as particularly unfair, given the moral compass defined by Big AI?
The situation strikes me as morally ambiguous. The resellers are: 1) selling Anthropic's products at a 95% discount and redirecting Anthropic's own customers to themselves. A customer is far less inclined to buy directly from Anthropic when a reseller is offering an identical product for 10x less. This situation is highly similar to internet piracy. 2) keeping the token logs from Anthropic's products and selling them…
For my part, I'll happily disclose that I have an axe to grind. I think the major AI labs are an aggressive form of a cancer that's been ravaging our society. I want to see them fail, of course -- but more than that, I want to see the public develop an immune response to this.
I just can't wrap my head around why someone would expend so much effort speaking up on their behalf. They have, after all, highly compensated PR people doing that for them!
Re: The Kimi K3 Moment
#499Earlier quoted context omitted.
The main story is what isn't being talked about. Chinese labs exfiltrated trillions of tokens of high-quality output from Anthropic and OpenAI, through proxies and heavily discounted token resellers, which they distilled and used for training data for their own models. Instead of spending 12-18 months building their own robust harnesses and painstakingly creating quality training data (which is what Anthropic and Ope…
>through proxies and heavily discounted token resellers Could you explain a little more about how this works? Are you saying that the Chinese run or have backdoored something like OpenRouter?
Chinese resellers acquire hundreds of Claude Max 5x accounts and set up a custom proxy server. Customers point their ANTHROPIC_API_KEY at that proxy, and requests are routed to Anthropic through one of those hundreds of accounts. Because one $200 Claude Max 5x account gets the equivalent of ~$2000 in of API credits, these resellers can resell Anthropic tokens at a massive discount, undercutting official API prices by more than 90%.
To cut costs even further, these accounts are funded using educational discounts, startup credits, or stolen credit cards.
The resellers log all data traveling through their proxy networks, which they then resell to Chinese labs as high-quality training data for significant profit. https://x.com/xkajon/status/2050445443889525235
The resellers also loan these proxy networks to Chinese labs, allowing them to can run distillation attacks on Anthropic, while blending in with regular user traffic. https://www.anthropic.com/news/detecting-and-preventing-dist...
This is a widespread tactic, there's hundreds of proxy resellers operating. Some even offer enterprise SLAs.
Re: The Kimi K3 Moment
#500Earlier quoted context omitted.
There's just really no incentive all they really want is just to train on that data to improve performance which in turn actually benefits your usecase since it becomes trained on that data and made available back to you. American labs take that data anyway and store it for years to possibly report you for misuse in the future for whatever reason they want. For example: you're very critical of X so they pull up your…
really weird that they would download every SF-86 file the government had and Equifax credit records of every American then.