Live data from Hacker News

The Kimi K3 Moment

stephen.bochinski.dev

91–100 of 644 posts

Re: The Kimi K3 Moment

#91
post #26

> The prices are nowhere near each other. K3’s API runs $3 per million input tokens and $15 per million output. Claude’s top model costs $10 and $50 for the same units. And this is the point where your internal compiler should have started shouting 'Type Error' Notice the trick here? > Then there’s the fine print. Claude couldn’t sustain Fable access on the twenty dollar plan, so they turned it off, and the plan quie…

People pay a premium for the best.

Re: The Kimi K3 Moment

#92

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

assume you are a "second class lab" and you are in fact making progress by distilling the results of the frontier labs' efforts. what is the end game for this strategy? if the frontier labs shut down, or stop releasing to the public, and there's noting left to distill, how will you progress?

This line of thinking makes no sense because it assumes that labs that distill from frontier models are doing nothing else. It's the classic "the Chinese can only copy" mentality, and it's going to end poorly for American companies.

I'm pretty sure that all labs are distilling each others' LLMs, maybe apart from Anthropic and OpenAI. It would be stupid not to do it, because it's cheap and effective. But that's not the only thing they're doing. If you think K3 and GLM-5.2 got this good only from distilling frontier models, you're not paying attention to Chinese labs' publications.

Re: The Kimi K3 Moment

#93
post #14

Earlier quoted context omitted.

How do we know that Chinese models would not do the same? What makes you so sure that China is less likely to abuse my data?

There's just really no incentive all they really want is just to train on that data to improve performance which in turn actually benefits your usecase since it becomes trained on that data and made available back to you. American labs take that data anyway and store it for years to possibly report you for misuse in the future for whatever reason they want. For example: you're very critical of X so they pull up your…

really weird that they would download every SF-86 file the government had and Equifax credit records of every American then.

Re: The Kimi K3 Moment

#94

This was always where this was heading, but we got here much faster than expected. Once western governments declare it to be a "national security" risk for citizens to have access to open-weight frontier models, and once they classify using these models as acts of terrorism, what will that world be like? Will using Kimi K3 come to be like how napster was in the olden days? Everybody knew it was technically illegal, b…

even 8x rtx pro 6000 is only 768GB of VRAM. IDK how anyone is going to run k3

Free server racks for everyone when the bubble bursts!

Re: The Kimi K3 Moment

#95

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

> Distillation “attacks” are not attacks. If "distillation attacks" happen, we have to conclude there is some value add in what model labs do. Regardless of how we feel about using existing human knowledge in the way they currently do, it's simply impractical to infer that everything that happens downstream of LLMs can not be an attack on some IP because of it. So both things can be true: a) People infringe on Anthro…

anthropic model output is not their IP

that would be existential doom for them because then they have a case to claim ownership of their users' codebases

no corporation would sign off on that

Re: The Kimi K3 Moment

#96

Even in this very thread the feedback on Kimi's actual efficacy is debated. I personally feel its worse than both Fable and 5.6 Sol, but I feel like the conversation isn't really about whether its good or not, but a backlash against the U.S governments foray into regulation. So I think people _want_ it to be superior out of anger/frustration with the current situation.

Agree completely.

Re: The Kimi K3 Moment

#98

Regardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a ch…

> Distillation “attacks” are not attacks. If "distillation attacks" happen, we have to conclude there is some value add in what model labs do. Regardless of how we feel about using existing human knowledge in the way they currently do, it's simply impractical to infer that everything that happens downstream of LLMs can not be an attack on some IP because of it. So both things can be true: a) People infringe on Anthro…

> People infringe on Anthropics IP

Unless someone literally stole the weights somehow (which is not out of the question, I doubt either oAI/Anthropic have the capabilities to prevent a state-level actor getting those weights), distillation from generations is not infringement on anyone's IP nor is it stealing nor is it an attack. It can't be. As long as you pay for tokens you get to do whatever you want with them. Someone saying you can't doesn't mean it's an attack or their IP or whatever. They either sell the tokens or not. They can decide to not sell them to anyone, but again that's not stealing.

And their ToS are a joke. Imagine how people would react if MS had ToS saying that you can't use MS software to develop solutions that compete with MS. They'd be laughed out of the room. Somehow it's ok for token sellers to decide what you do with the tokens? Why? If you pay for something you get to do whatever you want with that output. Train, distill, whatever.

Re: The Kimi K3 Moment

#100

I never truly understood what the intended business model around LLMs was. Get them widespread through cheap pricing and then jacking it up? Being the only ones that had a viable product so to get the ability to extract as much value as you want from AI? I don't understand how a product that: - is interfaced with and is deeply linked to natural language, so everything you produce (sessions, history, etc) is in Markdo…

> I never truly understood what the intended business model around LLMs was. A closely related question is “what do the American labs need to do in order to justify their enormous market valuations?” It seems like the answer cannot possibly be “gradually improve model capability while figuring out how to better monetize inference.” The valuations are just way too high for that to be sufficient. Surely the answer has…

> I never truly understood what the intended business model around LLMs was

What appeared initially to be a huge innovation was later easily duplicated by many. There are no platform-lockins or network effects. Switching costs for users are zero, and there are low barriers to entry, with vast numbers of models to choose from and more appearing every day. As a business a token will be a commodity like an electron. Doesnt matter who produces it, or how (solar, wind, coal, nuclear etc) as long as it powers my toaster.

Post reply on HN