Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

251–260 of 644 posts

Re: The Kimi K3 Moment

#251

Earlier quoted context omitted.

> empirically, it appears that distillation of a more advanced model is a required first step I see no evidence for that. > if this were not the case, then we would be observing chinese models that far surpass frontier models It's pretty clear that the primary reason for the difference is budget and compute availability. Chinese labs have at least an order of magnitude less money than Anthropic and OpenAI. > what hap…

https://www.anthropic.com/news/detecting-and-preventing-dist... Moonshot AI Scale: Over 3.4 million exchanges The operation targeted: Agentic reasoning and tool use Coding and data analysis Computer-use agent development Computer vision Moonshot (Kimi models) employed hundreds of fraudulent accounts spanning multiple access pathways. Varied account types made the campaign harder to detect as a coordinated operation.…

I'm assuming you posted that as evidence for the claim that "empirically, it appears that distillation of a more advanced model is a required first step", but I don't think it is. It's just evidence that Moonshot distills Anthropic's models, which, yes, they do.

Re: The Kimi K3 Moment

#252

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

> Distillation “attacks” are not attacks. If "distillation attacks" happen, we have to conclude there is some value add in what model labs do. Regardless of how we feel about using existing human knowledge in the way they currently do, it's simply impractical to infer that everything that happens downstream of LLMs can not be an attack on some IP because of it. So both things can be true: a) People infringe on Anthro…

> People infringe on Anthropics IP

Anthropic’s model outputs contain no IP. This is actually a simple legal proposition (rare in this field!) that derives from the fact that only specific classes of IP exist: copyrights, patents, trade secrets, and trademarks. Examining each, it is clear that API outputs do not qualify. Anthropic disclaims copyright in outputs; the outputs are not patented; the outputs are not secret (a prerequisite to having trade secrets); and trademarks are irrelevant in concept.

Re: The Kimi K3 Moment

#253

Pricing is actually far cheaper than that. There's two tiers of pricing: Chinese and US. If you sign up with non-Chinese phone number, you're bucketed into US, you get US prices, can pay only in USD and with American credit card network. Chinese prices are about 9x cheaper than the US prices, which are already far cheaper than Claude or other American provider. If you can somehow get hold of a Chinese phone number, k…

Most of this hand-wringing on price will go away.

My assumption is that Anthropic, OpenAI, Kimi, etc all have a similar cost structure when serving models. The same size model roughly generates the same GPU usage whether you’re American or Chinese. I’d also guess that the model sizes across all SOTA models is similar, we just only see data for open models. The difference is most likely that American companies simply charge more because they have the dominant market position.

Remember not too long ago when Anthropic was charging $75/mt for Opus? Now that many models are in “opus tier”, their pricing is $25 - higher than competitors but close. The newest Kimi is $15. 40% lower to forgo “made in America” with American enterprise support staff is not crazy. Compare AWS to Hetzner or any other flagship enterprise service to the foreign and discount option. I assume that over time, we’ll see the commodification of models reducing prices even towards the raw GPU costs.

Re: The Kimi K3 Moment

#254

Earlier quoted context omitted.

the question was: what is the endgame for the stated "second class labs" strategy of distilling their frontier competitors then undercutting them on price?

Yes yes, we all understand the game-theoretic race-to-the-bottom you're describing here. Somehow despite linux being FOSS it still powers most of the important computing in the world. Can you explain how that works despite it being free? Once you understand that case I think you'll understand the game-theory behind how large projects can exist in the absence of traditional IP protection.

the obvious difference is the massive scale of data and compute required to develop and evolve these models, and the costs they impose on those building them.

Re: The Kimi K3 Moment

#255

Earlier quoted context omitted.

>People infringe on Anthropics IP No. Authors do not infringe on IP when they read another's book, nor should the lumber company be able to dictate how I use planks and if I can resell them if i'm done with them. You're framing it as if the added value of the author or lumber company, awards them consideration when somebody uses the products to create more value. IP law was always a big mess, and these questions cros…

It's more simple: They infringe on the IP by way of violating the ToS. If you violate ToS and the company suffers financial harm, they usually can (usually) sue you in civil court for damages.

That’s not what “IP” means. You’re describing breach of contract.

Re: The Kimi K3 Moment

#256

Pricing is actually far cheaper than that. There's two tiers of pricing: Chinese and US. If you sign up with non-Chinese phone number, you're bucketed into US, you get US prices, can pay only in USD and with American credit card network. Chinese prices are about 9x cheaper than the US prices, which are already far cheaper than Claude or other American provider. If you can somehow get hold of a Chinese phone number, k…

It’s 100 yuan per million output tokens in China. That’s $14.7 USD - not “far cheaper”.

Re: The Kimi K3 Moment

#257
post #178

Earlier quoted context omitted.

Eh, I think you've done a pretty good job summarizing a collection of settlements with a few narrow bench rulings for seasoning. I'm not sure I follow you to it being a coherent legal theory. Buying a book in a bookstore is sure legal, and excerpting from it for e.g. literary criticism is pretty settled. Downloading every torrent of all e-books ever is pretty clearly illegal (or at least it fuckin would be if I did i…

> Downloading every torrent of all e-books ever is pretty clearly illegal (or at least it fuckin would be if I did it). Pretty sure like, multiple labs have been popped for that though. Oh it is, and at least anthropic has paid $1.5 billion and deleted there torrented copies and not released any models derived from them as a consequence. The thing is it turns out to be not that expensive to just buy a copy of every b…

> and deleted there torrented copies and not released any models derived from them as a consequence.

I have a bridge to sell you

Re: The Kimi K3 Moment

#258
post #8

I tried Kimi K3 on a task I've done with every other model I use regularly ( https://swelljoe.com/post/i-let-every-agent-implement-its-ow... ) and found it chewed a lot longer on the problem and ate up almost the entirety of a 5 hour usage limit on their $19 plan. I only have the $20 plan from OpenAI and the same task, with a lot of the same implementation details as Kimi Code, only took a few minutes and consumed al…

We really need to stop using $/M tokens as the pricing benchmark. I've found that the number of tokens used tends to be a bigger factor than the listed per token price. The cost per task vs. intelligence curve is really what you care about, and in my estimation Chinese models are just not there. They are focused on benchmaxing and getting the highest raw score they can, rather than efficiency.

Re: The Kimi K3 Moment

#259

Earlier quoted context omitted.

https://www.anthropic.com/news/detecting-and-preventing-dist... Moonshot AI Scale: Over 3.4 million exchanges The operation targeted: Agentic reasoning and tool use Coding and data analysis Computer-use agent development Computer vision Moonshot (Kimi models) employed hundreds of fraudulent accounts spanning multiple access pathways. Varied account types made the campaign harder to detect as a coordinated operation.…

I'm assuming you posted that as evidence for the claim that "empirically, it appears that distillation of a more advanced model is a required first step", but I don't think it is. It's just evidence that Moonshot distills Anthropic's models, which, yes, they do.

it is not a required first step for training a model, sure. but that's not what i claimed. what i claimed is that is how they are so significantly _reducing the cost_ of training one! how else do you think they are doing it?

Re: The Kimi K3 Moment

#260
post #226

Earlier quoted context omitted.

> to someone with rights to serve the contents Now THAT'S doing some heavy lifting lmao. The vast, vast, VAST majority of the original datasets were from pirated books and the like. Also, arguably a robots.txt is the exact mechanism to follow to do the mass GET-ing, yet the AI cos choose time and time and time again to simply ignore it and be as abusive as they possibly fucking can

> The vast, vast, VAST majority of the original datasets were from pirated books and the like And there's been significant legal consequences as a result > Also, arguably a robots.txt is the exact mechanism to follow to do the mass GET-ing You're free to argue this of course, but the courts have largely rejected it already pre LLMs. See for example hiQ Labs v. LinkedIn

Yes, anthropic and openai have really been brought to their knees and ipos cancelled because of the legal consequences of obtaining their training data.
Post reply on HN