Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

181–190 of 644 posts

Re: The Kimi K3 Moment

#182
post #146
post #29

Earlier quoted context omitted.

The visa that would correlate to this is the O-1 visa 20k O-1 visas were issued last FY which was mostly under the Trump admin, up from 19.5k the previous FY under the Biden admin

The O-1 has also been abused for a long time, basically any software engineer kid who gets into Y Combinator has been getting an O-1

This is essentially the point of the visa, it feels wrong especially as YC drops standards and increases cohort sizes, but the same power laws that keep them winning also apply here in maximizing economic value of each O-1 approval

Basically of all visas O-1 is virtually guaranteed to have highly positive economic value

Re: The Kimi K3 Moment

#183

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

>Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models So why didnt we have these LLMs in 2005?

Because the transformer architecture that enabled modern LLMs wasn't invented until 2017[1]?

1: That's the "T" in GPT fyi, even though Google is the author of the research paper that changed everything

Re: The Kimi K3 Moment

#184
post #164

Earlier quoted context omitted.

Right. Therefor it's not simply a derivative of information. The hardware is required to build the model. Software as well. The model uses information, it is not "distilled" from it. "Distillation" literally means to separate and take some components out of something. You can distill how a model works from a model. You cant distill a model from information because the information does not contain the model. People ar…

Information is information. Why is some information considered different than others in your estimation?

> Information is information

Not really, what the information actually is, matters a great deal. It's harder to get good results going from "nothing > model+weights" than "nothing + traces from known good sessions of other good model > model+weights", this is what the "distillation" part is referring to. If "information is information", you wouldn't even need to separate good from bad sessions while doing the training, which leads to somewhat obvious results if you don't.

Re: The Kimi K3 Moment

#185

Earlier quoted context omitted.

Right. Therefor it's not simply a derivative of information. The hardware is required to build the model. Software as well. The model uses information, it is not "distilled" from it. "Distillation" literally means to separate and take some components out of something. You can distill how a model works from a model. You cant distill a model from information because the information does not contain the model. People ar…

That argument is moot as distillation also requires a lot of hardware and software, if copying models was as easy as that, we would have hundreds of competing models.

No. Building models and distilling models both require the hardware and software. It doesn't mean building models is distillation.

Re: The Kimi K3 Moment

#186
The dumb efforts by the US AI industry to use fear mongering for regulatory capture will hand dominance to China and others.

In a few years there will be Mythos level open weight models hosted by the lowest bidder anyway.

Re: The Kimi K3 Moment

#187
post #164

Earlier quoted context omitted.

Right. Therefor it's not simply a derivative of information. The hardware is required to build the model. Software as well. The model uses information, it is not "distilled" from it. "Distillation" literally means to separate and take some components out of something. You can distill how a model works from a model. You cant distill a model from information because the information does not contain the model. People ar…

Information is information. Why is some information considered different than others in your estimation?

Can you be more specific? I have no idea what you are trying to say.

To succinctly restate my point, you cannot distill a model from information because the model is not contained within that information. You can distill a model from another model.

Re: The Kimi K3 Moment

#188

This was always where this was heading, but we got here much faster than expected. Once western governments declare it to be a "national security" risk for citizens to have access to open-weight frontier models, and once they classify using these models as acts of terrorism, what will that world be like? Will using Kimi K3 come to be like how napster was in the olden days? Everybody knew it was technically illegal, b…

The disappearance of high ram Mac studio rigs is probably just a coincidence, right? :/

Re: The Kimi K3 Moment

#189

Earlier quoted context omitted.

i never assumed that, and i do keep up with the publications. i'm also not saying it's a dumb thing to do! what i am saying is that empirically, it appears that distillation of a more advanced model is a required first step for them to train a borderline competitive, cheaper model. in effect, their training is subsidized by the frontier labs. if this were not the case, then we would be observing chinese models that f…

> empirically, it appears that distillation of a more advanced model is a required first step I see no evidence for that. > if this were not the case, then we would be observing chinese models that far surpass frontier models It's pretty clear that the primary reason for the difference is budget and compute availability. Chinese labs have at least an order of magnitude less money than Anthropic and OpenAI. > what hap…

https://www.anthropic.com/news/detecting-and-preventing-dist...

Moonshot AI Scale: Over 3.4 million exchanges

The operation targeted:

Agentic reasoning and tool use Coding and data analysis Computer-use agent development Computer vision Moonshot (Kimi models) employed hundreds of fraudulent accounts spanning multiple access pathways. Varied account types made the campaign harder to detect as a coordinated operation. We attributed the campaign through request metadata, which matched the public profiles of senior Moonshot staff. In a later phase, Moonshot used a more targeted approach, attempting to extract and reconstruct Claude’s reasoning traces.

Re: The Kimi K3 Moment

#190

Earlier quoted context omitted.

>Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models So why didnt we have these LLMs in 2005?

Because the transformer architecture that enabled modern LLMs wasn't invented until 2017[1]? 1: That's the "T" in GPT fyi, even though Google is the author of the research paper that changed everything

Right. So we had enough information to train LLMs but not the technology to build it.

So the initial models arent just distilled from information. We’ve always had the information.

Post reply on HN