Earlier quoted context omitted.
In this world too, the punishment was, ballpark, 100x direct damages.
> In this world too, the punishment was, ballpark, 100x direct damages. Which ballpark are you playing in to come to a number like that? I have friends whom are authors, and they've certainly not seen a penny from Anthropic. I somehow doubt they're the only ones out there that haven't been compensated for having their work blatantly stolen. In a just world we'd be ensuring Anthropic was destroyed as a result of their…
The Kimi K3 Moment
581–590 of 644 posts
Re: The Kimi K3 Moment
#582Earlier quoted context omitted.
The artificialanalysis cost per task chart has DeepSeek as the clear winner and Fable as the clear loser. But I would still pick Fable for some tasks, so that also can't be all there is to it. But I agree that price per token figure is not great. It seems even the tokens per character can vary between models, so it's basically useless.
Wow, you weren't kidding. I looked at their chart, and the cost-per-task for Fable is more than double Sol's. And DeepSeek absolutely stomps. Four cents per-task vs Sol's $1 and Fable's $3. I might need to check out DeepSeek more. I had no idea the difference was this obscene. Makes me wonder if something's off with the benchmark. A 70x cost reduction vs. Fable seems too good to be true.
GPT 5.5 Pro was ~230x at almost $23 per task.
DeepSeek is my go-to when I need an API, and local Gemma 4 won't do because it's either too slow or not capable enough. DeepSeek isn't at the frontier but it's good enough for a lot of things, very cheap, and quite fast. Flash is even faster and cheaper, and still better than anything I can host locally.
Re: The Kimi K3 Moment
#583Earlier quoted context omitted.
Given how OpenAI got rid of their 5-hour limits and reset weekly limits so often, is Kimi really undercutting them on effective price?
The 5 hour limits are coming back soon right? I thought that was temporary
Re: The Kimi K3 Moment
#584Earlier quoted context omitted.
> In this world too, the punishment was, ballpark, 100x direct damages. Which ballpark are you playing in to come to a number like that? I have friends whom are authors, and they've certainly not seen a penny from Anthropic. I somehow doubt they're the only ones out there that haven't been compensated for having their work blatantly stolen. In a just world we'd be ensuring Anthropic was destroyed as a result of their…
Anthropic paid (or is paying, not sure the schedule for class members receiving payouts) 3000 per work they were found to have infringed - 30 is a fair approximation of the cost to buy a book.
Re: The Kimi K3 Moment
#585Earlier quoted context omitted.
So the efficient market hypothesis is wrong?
How is the efficient market hypothesis applicable here?
Re: The Kimi K3 Moment
#586Earlier quoted context omitted.
Kimi calling itself claude means nothing. During pre-training, when the model learns to "simulate" the internet text, it will naturally be fed with a bunch of data about Claude and ChatGPT. With the amount of LLM outputs on the internet today, it is not surprising at all that a model would naturally call itself Claude or ChatGPT. You can mitigate that in post-training (or actually in pre-training as well) by training…
Sure, but then Qwen should leak that too, and it doesn't. K3 calls itself Claude 7 out of 48 times, Qwen does it 0 out of 48, and the only other model to identify itself as Claude is DeepSeek. and DeepSeek is alleged to also distill from Claude data anyway. So this isn't something every model absorbed from the same web text. And you skipped over the strongest datapoint that K3 is distilled: K3 reproduces Claude's pub…
This doesn't imply that Kimi's reasoning capabilities or coding ability has anything to do with Anthropic (which is the most valuable part), or the model's strength comes from distilling Anthropic.
Considering by this benchmark, Opus sounds extremely similar to Fable, this isn't really evidence of a ton of Fable output being trained on (and even then, most likely for superficial stuff, like style)
I also have a suspicion that in original Chinese, these models don't sound anything like Anthropic's ones, though I have no proof of that.
Re: The Kimi K3 Moment
#587Earlier quoted context omitted.
> Kimi K3 reproducibly identifies itself as Claude It could also be have been trained from collected response datasets. Claude got caught several time responding it was ChatGPT or even Deepseek and I don't think Anthropic has been distealling DeepSeek. > This behavior is exactly what you'd expect from a model distilled from Claude. The opposite actually. If they wanted to distill Claude without getting caught they co…
> distealling Apt typo. Though I am of the opinion that distilling is no different than how extant frontier LLMs have also been trained on other people's data, I could actually see the word distealling becoming useful in discussion.
While still not okay, I suspect the latter is what gets stolen by Chinese distillation (and some evidence suggest this happens the other way round with US models talking in Chinese)
Re: The Kimi K3 Moment
#588Earlier quoted context omitted.
Anthropic paid (or is paying, not sure the schedule for class members receiving payouts) 3000 per work they were found to have infringed - 30 is a fair approximation of the cost to buy a book.
And the typical book only sees 1 copy sold?
Re: The Kimi K3 Moment
#589Earlier quoted context omitted.
The main story is what isn't being talked about. Chinese labs exfiltrated trillions of tokens of high-quality output from Anthropic and OpenAI, through proxies and heavily discounted token resellers, which they distilled and used for training data for their own models. Instead of spending 12-18 months building their own robust harnesses and painstakingly creating quality training data (which is what Anthropic and Ope…
assuming the k3 model weights do indeed get published, if your model of the world is "achieving RSI is beneficial and K3 has done so," this feels structurally different from ordinary industrial espionage, because the knowledge has enriched the commons more like silk than capacitors if, again, your model is that RSI will be beneficial, why wouldn't making it available to all unlock more benefit globally than not doing…
Re: The Kimi K3 Moment
#590Earlier quoted context omitted.
I am still shocked Spain/Italy and USA are considered 'first world' countries. We are not in 70s or even 90s anymore. I've been to China in 2011 also thinking I am visiting some huge village but...that was the most futuristic trip I ever had. I was surprised by the penetration level of the mobile devices - everything had a QR code, you could buy/sell/send money, pay services all with a single tap on a phone.
I mean, you have those kind of luxuries even in the poorest of countries in Asia, it’s just that there’s still a huge discrepancy between rich and poor, city vs countryside. It’s not difficult to find areas in all these countries that are significantly less developed than Spain/Portugal’s underdeveloped areas. It’s just not as black and white as you seem to suggest. (I come from EU but have been living in various cou…
But that's because in a market economy, you can't just give away things, and those poor people don't produce much of value by capitalist standards, so the can't pay for those things.
I'm sure the CCP is working on the problem of how to give these people those things without crashing the economy.