Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

201–210 of 644 posts

Re: The Kimi K3 Moment

#201
It was all distillation up to this point anyway. And I agree with what Suhail said on twitter: "Make the margins next to zero for all these AI models. It was trained on humanity's data, it should be gift to ourselves. Doing so will save us from a few in control of our species."

Re: The Kimi K3 Moment

#202

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

> Distillation “attacks” are not attacks. If "distillation attacks" happen, we have to conclude there is some value add in what model labs do. Regardless of how we feel about using existing human knowledge in the way they currently do, it's simply impractical to infer that everything that happens downstream of LLMs can not be an attack on some IP because of it. So both things can be true: a) People infringe on Anthro…

Anthropic’s IP is basically null and void for how they created it. And they might not want to try and challenge this in court, considering how they had to settle for using text books they had no right to use

Re: The Kimi K3 Moment

#203

Earlier quoted context omitted.

Us models didnt pay for licenses too

That is incorrect. Anthropic paid $1.5 billion in compensation to copyright holders for use of their content in training data. OpenAI pays hundreds of millions per year across 150+ licensing deals for access to copyrighted data. Meta and Alphabet have similar arrangements. Under the settlement, Anthropic was forced to delete the pirated data they were training on. Chinese labs can still train on pirated data. I doubt…

Let’s not forget that Anthropic only paid that to settle a class action lawsuit.

Re: The Kimi K3 Moment

#204

Earlier quoted context omitted.

>People infringe on Anthropics IP No. Authors do not infringe on IP when they read another's book, nor should the lumber company be able to dictate how I use planks and if I can resell them if i'm done with them. You're framing it as if the added value of the author or lumber company, awards them consideration when somebody uses the products to create more value. IP law was always a big mess, and these questions cros…

It's more simple: They infringe on the IP by way of violating the ToS. If you violate ToS and the company suffers financial harm, they usually can (usually) sue you in civil court for damages.

You can't violate ToS you never agreed to. If I use pirate Claude through a third-party reseller, I have entered no agreement with Anthropic.

Re: The Kimi K3 Moment

#205
I think it's the opposite.

Kimi K3 has 2.8 trillion parameters. We don't know the number of parameters of ChatGPT 5.6 or Opus 4.8, but it's probably in the same region. Fable/Mythos are rumored to be around 10 trillion.

So, K3 is directly comparable with ChatGPT 5.6 and Opus 4.8, and the price is not so much lower:

K3: $3/$15 per 1 Mtok input/output ChatGPT 5.6 Sol: $5/$30 Opus 4.8: $5/$25

This is not a watershed moment. It's a competitor converging to the same capability and trying to undercut your prices, but not by a lot.

As for the open weights? For now, Kimi K3's weights are closed, and I don't expect the situation would change.

Re: The Kimi K3 Moment

#206
post #8

I tried Kimi K3 on a task I've done with every other model I use regularly ( https://swelljoe.com/post/i-let-every-agent-implement-its-ow... ) and found it chewed a lot longer on the problem and ate up almost the entirety of a 5 hour usage limit on their $19 plan. I only have the $20 plan from OpenAI and the same task, with a lot of the same implementation details as Kimi Code, only took a few minutes and consumed al…

> Subscription usage limits are hard to measure as none of the providers tell you directly what it means in terms of tokens or anything else you can easily compare AI subscription pricing is so goofy. You get some amount of usage that varies by models, is measured by opaque token usage, driven by how many tokens the (usually) vendor-provided interface (or model itself) wants to use. Then your usage is limited by time…

AI subscription pricing was fine when it was $100/month for some opaque 5 hour token budget I don't think I ever used, not even that one day where I coded for 14 hours non-stop using Fable. But like most people with low token usage, I had a human in the loop and and I didn't use workflows with swarms of agents.

Now, of course, the plan is to remove Fable from the subscription. To paraphrase Darth Vader, they have altered the deal. Pray they do not alter it further.

Re: The Kimi K3 Moment

#207
post #24

The current administration's immigration policy isn't helping. This wouldn't have happened 10 years ago because the US was this city on the hill that everyone wanted to immigrate to. Talented Asian researchers would have immigrated to the US and China would be deprived of talent.

China never allow US AI in China, so they HAVE to build Chinese equivalents...

US immigration policy isn't a big factor.

China's got 1.8B people. If you don't think they've got the talent to pull this off, even if a lot of it leaves to live elsewhere, you're naive.

No one uses Baidu, but they built their own Google, and it's good.

They built their own Facebooks and Instagrams.

The US isn't the only place in the world where people can build software...

Re: The Kimi K3 Moment

#208

Earlier quoted context omitted.

Us models didnt pay for licenses too

That is incorrect. Anthropic paid $1.5 billion in compensation to copyright holders for use of their content in training data. OpenAI pays hundreds of millions per year across 150+ licensing deals for access to copyrighted data. Meta and Alphabet have similar arrangements. Under the settlement, Anthropic was forced to delete the pirated data they were training on. Chinese labs can still train on pirated data. I doubt…

Anthropic paid $1.5 billion in compensation to copyright holders for use of their content in training data.

The payment was for illegally downloading copyrighted material, not training. Training was explicitly ruled to be fair use.

Re: The Kimi K3 Moment

#209

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

> Distillation “attacks” are not attacks. If "distillation attacks" happen, we have to conclude there is some value add in what model labs do. Regardless of how we feel about using existing human knowledge in the way they currently do, it's simply impractical to infer that everything that happens downstream of LLMs can not be an attack on some IP because of it. So both things can be true: a) People infringe on Anthro…

Considering they were the original infringers, I don't know how anyone can expect tears to be shed here. The best we can hope for is for all these cancerous - and they really are the definition of a cancer - money burning entities to all fall apart to distillation attacks like these.

Re: The Kimi K3 Moment

#210
post #164

Earlier quoted context omitted.

Information is information. Why is some information considered different than others in your estimation?

Can you be more specific? I have no idea what you are trying to say. To succinctly restate my point, you cannot distill a model from information because the model is not contained within that information. You can distill a model from another model.

Their point is that "training" and "distillation" are essentially the same. The difference between the words is whether the source material is output from another model, vs being some original text.
Post reply on HN