Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

81–90 of 644 posts

Re: The Kimi K3 Moment

#81

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

> Distillation “attacks” are not attacks. If "distillation attacks" happen, we have to conclude there is some value add in what model labs do. Regardless of how we feel about using existing human knowledge in the way they currently do, it's simply impractical to infer that everything that happens downstream of LLMs can not be an attack on some IP because of it. So both things can be true: a) People infringe on Anthro…

I'm pretty sure that LLM output is not intellectual property. Nobody owns it, and it can't even be copyrighted. So using output from Anthropic's LLMs in ways Anthropic does not condone is not IP infringement.

Re: The Kimi K3 Moment

#82

This was always where this was heading, but we got here much faster than expected. Once western governments declare it to be a "national security" risk for citizens to have access to open-weight frontier models, and once they classify using these models as acts of terrorism, what will that world be like? Will using Kimi K3 come to be like how napster was in the olden days? Everybody knew it was technically illegal, b…

even 8x rtx pro 6000 is only 768GB of VRAM. IDK how anyone is going to run k3

This is why God gave us 1.58-bit ternary quants?

Re: The Kimi K3 Moment

#84
post #80

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

or the application layer - which will capture majority of the value. yeah hardware companies make for nice stories or green numbers on Wall Street - but value will be captured by application layer. look at history.

That’s true up until the point where you can ask the hardware you made to make its own application layer.

Re: The Kimi K3 Moment

#85

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

Almost all markets depend on some form of regulation whether its as simple as "leave everyone alone but no stealing" or "every participant has to source every object through mountains of red tape." Thus far the US has not really chosen to go the Chinese rare-earth method yet. The problem with distillation attacks is the end result is everyone who is not doing them is going to deal with some kind of regulation whether…

Given these models could not have been trained in the first place if they had to license every line of random fan fiction on the internet, I think distillation also being fair game is a tradeoff everyone should be willing to take (unless they want to decelerate, but that's a different conversation).

Re: The Kimi K3 Moment

#86
post #42

Earlier quoted context omitted.

Yes, this (imo) is a clear result of benchmaxxing. You can get a much better score on most "intelligence" benchmarks by massively over-saturating reasoning. This looks good on those, but for actual daily usage makes the models much less effective: I don't want a model I use for coding to burn a bunch of reasoning (read: time) on trivial tasks.

I strongly suspect the flip side is that in the future it enables you to train smarter models by "distilling" the end result of the super duper heavily thinking models.

But these models already distill the smarter American ones ;)

Re: The Kimi K3 Moment

#87

Earlier quoted context omitted.

> Distillation “attacks” are not attacks. If "distillation attacks" happen, we have to conclude there is some value add in what model labs do. Regardless of how we feel about using existing human knowledge in the way they currently do, it's simply impractical to infer that everything that happens downstream of LLMs can not be an attack on some IP because of it. So both things can be true: a) People infringe on Anthro…

>People infringe on Anthropics IP No. Authors do not infringe on IP when they read another's book, nor should the lumber company be able to dictate how I use planks and if I can resell them if i'm done with them. You're framing it as if the added value of the author or lumber company, awards them consideration when somebody uses the products to create more value. IP law was always a big mess, and these questions cros…

There are some quite interesting legal implications here. If Anthropic has IP over output produced by agents, do they somehow have legal rights to code and documents produced by such agents?

This would demolish agent usage by corporations.

Re: The Kimi K3 Moment

#88

Earlier quoted context omitted.

In public with budgets that don't risk destroying the American economy presumably. Yes it may be slower.

> with budgets and what will fund these budgets exactly? inference is cheap, distillation is cheap, training is what's expensive.

Presumably the US military / NSA.

Re: The Kimi K3 Moment

#89

I never truly understood what the intended business model around LLMs was. Get them widespread through cheap pricing and then jacking it up? Being the only ones that had a viable product so to get the ability to extract as much value as you want from AI? I don't understand how a product that: - is interfaced with and is deeply linked to natural language, so everything you produce (sessions, history, etc) is in Markdo…

[deleted]

Re: The Kimi K3 Moment

#90
Even in this very thread the feedback on Kimi's actual efficacy is debated. I personally feel its worse than both Fable and 5.6 Sol, but I feel like the conversation isn't really about whether its good or not, but a backlash against the U.S governments foray into regulation. So I think people _want_ it to be superior out of anger/frustration with the current situation.
Post reply on HN