Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

141–150 of 644 posts

Re: The Kimi K3 Moment

#141
I can see the economics of open vs. frontier models turning out similarly to pharmaceuticals, where generic drugs cost a fraction what the name brands do and Americans end up paying the highest prices in the world partly as a consequence of propping up drug discovery research.

Re: The Kimi K3 Moment

#142
post #134

Earlier quoted context omitted.

Contract law is never going to prevent this.

Why not? Seems like a perfectly normal contact term to me. Or do you just mean that US courts don't have enough teeth to prevent Chinese companies from violating contracts? On that I agree.

I mean the latter, but more narrowly: China would never allow the United States to have a monopoly on machine intelligence if the only thing standing in the way of a domestic alternative was the Anthropic ToS. In general, I think that China is willing to agree on certain things relating to intellectual property. But not on this, it’s too big.

The US is already publicizing the way they are using Claude with Palantir for war gaming purposes. It’s a matter of national defense. Contract law has no meaning here.

Re: The Kimi K3 Moment

#143

Earlier quoted context omitted.

assume you are a "second class lab" and you are in fact making progress by distilling the results of the frontier labs' efforts. what is the end game for this strategy? if the frontier labs shut down, or stop releasing to the public, and there's noting left to distill, how will you progress?

This line of thinking makes no sense because it assumes that labs that distill from frontier models are doing nothing else. It's the classic "the Chinese can only copy" mentality, and it's going to end poorly for American companies. I'm pretty sure that all labs are distilling each others' LLMs, maybe apart from Anthropic and OpenAI. It would be stupid not to do it, because it's cheap and effective. But that's not th…

i never assumed that, and i do keep up with the publications. i'm also not saying it's a dumb thing to do! what i am saying is that empirically, it appears that distillation of a more advanced model is a required first step for them to train a borderline competitive, cheaper model. in effect, their training is subsidized by the frontier labs.

if this were not the case, then we would be observing chinese models that far surpass frontier models in capabilities, rather than "almost as good, but much cheaper", and we would be having a very different conversation. what happens to these efforts when the subsidy is cut off?

Re: The Kimi K3 Moment

#144

Earlier quoted context omitted.

>Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models So why didnt we have these LLMs in 2005?

Is this some form of rage bait? 2005 we hadn't the GPUs, we have today. There are other factors, but I think this is the big one. The mathematics of building an LLM are really old, we just hadn't the hardware to do the needed calculations.

Right. Therefor it's not simply a derivative of information. The hardware is required to build the model. Software as well. The model uses information, it is not "distilled" from it.

"Distillation" literally means to separate and take some components out of something. You can distill how a model works from a model. You cant distill a model from information because the information does not contain the model.

People are happy to conflate distilling with building because they dont like how the information was used. You distill how the model works from the model, and you build a model with information. Both could be morally good or bad but its not the same thing.

Re: The Kimi K3 Moment

#146
post #29
post #24

The current administration's immigration policy isn't helping. This wouldn't have happened 10 years ago because the US was this city on the hill that everyone wanted to immigrate to. Talented Asian researchers would have immigrated to the US and China would be deprived of talent.

The visa that would correlate to this is the O-1 visa 20k O-1 visas were issued last FY which was mostly under the Trump admin, up from 19.5k the previous FY under the Biden admin

The O-1 has also been abused for a long time, basically any software engineer kid who gets into Y Combinator has been getting an O-1

Re: The Kimi K3 Moment

#147

I never truly understood what the intended business model around LLMs was. Get them widespread through cheap pricing and then jacking it up? Being the only ones that had a viable product so to get the ability to extract as much value as you want from AI? I don't understand how a product that: - is interfaced with and is deeply linked to natural language, so everything you produce (sessions, history, etc) is in Markdo…

Same business model as always: build cool tech because it is cool and figure it out later.

Re: The Kimi K3 Moment

#148
post #136

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

Look how hard Anthropic is to even be able scroll back on your conversation, or look at the thinking tokens or subagents. They want to keep everyone coming back to the watering hole but never to learn how to dig a well.

Why is it hard to scroll?
Post reply on HN