Live data from Hacker News

Kimi-K3 Technical Report [pdf]

github.com

101–110 of 199 posts

Re: Kimi-K3 Technical Report [pdf]

#101
post #74

Earlier quoted context omitted.

> Kudos to Moonshot for truly being what OpenAI should have been. Kudos to Moonshot for making these weights available for download. Lets not fool ourselves and claim these are "open source" by any understanding of the concept though, there are usage restrictions (even if you download them) and also training data isn't clearly broken down either, nor it it actually using a FOSS license.

> Kudos to Moonshot for making these weights available for download. Lets not fool ourselves and claim these are "open source" by any understanding of the concept though, there are usage restrictions (even if you download them) and also training data isn't clearly broken down either, nor it it actually using a FOSS license. They will probably never release the training data because that represents a large part of the…

Exactly. I wish people would just stop saying "open source" regarding models. Even "open weight" is disingenuous. "self-hostable" would be more honest.

Re: Kimi-K3 Technical Report [pdf]

#102
post #78
post #67

Earlier quoted context omitted.

To me its clear that it is decel the only reason other labs can catch up is because the frontier labs can be distilled, and they siphon a % of the labs' revenue to reinvest into the next iteration full accel would mean nationalizing the big 2 labs and locking in manhattan project style until RSI (Edit: some great counterpoints in the replies. my view has definitely been changed!)

only if you only get your news from main stream business press and Big Lab propaganda channels There's no chance K3 is a distill of Fable, it came out way too soon after the limited fable release to be feasbile. If you look at all of the top ML conferences, chinese labs contribute way more to advances in ML than "Open"AI and Anthropic: https://www.reddit.com/r/TheMachineGod/comments/1pi4q7f/pape... This K3 release ju…

Appreciate this point. Back in grad school, I published a peer reviewed paper with all the source code and datasets used. Got heckled at a conference talk by staff from a commercial lab. They shouted, we figured this out five years ago, lol. Also our approach is still better. But they don’t release their source code or publish much, so no one knows what this approach is or if it’s actually better.

And years down the line, lots of other research labs used my code and cited my paper.

Re: Kimi-K3 Technical Report [pdf]

#103
post #85
post #61

Earlier quoted context omitted.

In what way was the license "janky"? Moonshot isn't a trillion dollar US behemoth. If these terms allow it to release near frontier models with open weights, so startups and anyone with enough hardware can use them, while charging only companies with more than $20 million in revenue or 100 million users, that seems like a reasonable trade off.

I use the term "janky" for any time someone releases something under a supposedly open source license that doesn't comply with the OSI definition. "Modified MIT" is the perfect example of that. I'm not saying it's unreasonable, or that you can't release under such a license - it's your software, use whatever license you like! I'll call it "janky" when you do.

So it's more like your personal definition, not the normal meaning of the word. Ok, because "janky" is informal slang also meaning poor quality, unreliable etc. I understand that you're using janky as shorthand for not OSI compliant. I just think that can blur two separate issues: whether a license qualifies as open source under the OSI definition, and whether the license itself is unreasonable/badly designed.

Re: Kimi-K3 Technical Report [pdf]

#104

Earlier quoted context omitted.

> And hire 2 or 3 dev ops to keep it running Not a devops but I'd say one full time is already too many.

Yes but zero is not enough and where do you get a fraction of a competent dev op from?

From the other dev-op work you’re doing.

Re: Kimi-K3 Technical Report [pdf]

#105
post #85

Earlier quoted context omitted.

I use the term "janky" for any time someone releases something under a supposedly open source license that doesn't comply with the OSI definition. "Modified MIT" is the perfect example of that. I'm not saying it's unreasonable, or that you can't release under such a license - it's your software, use whatever license you like! I'll call it "janky" when you do.

So it's more like your personal definition, not the normal meaning of the word. Ok, because "janky" is informal slang also meaning poor quality, unreliable etc. I understand that you're using janky as shorthand for not OSI compliant. I just think that can blur two separate issues: whether a license qualifies as open source under the OSI definition, and whether the license itself is unreasonable/badly designed.

I'm trying to make fetch happen.

Re: Kimi-K3 Technical Report [pdf]

#107

Earlier quoted context omitted.

If your worldview is “most of the progress is made by closed labs, then open labs fast-follow” (which isn’t implausible given the documented distillation of Fable), and further that open labs cannot make make meaningful progress vs the closed labs except by fast-following and that they won’t pick up the ability to make progress after the closed labs are gone, then driving closed labs out of business slows down overal…

I think it's pretty hard to hold that worldview: Anthropic couldn't ship a reasoning model until they copied DeepSeek R1's homework, and they've all copied DS-style super-sparse MoEs at this point too.

That’s a really good point. Folks really need to read the papers coming out of these Chinese labs. Every paper from the DeepSeek team has been a step change.

Re: Kimi-K3 Technical Report [pdf]

#108

Earlier quoted context omitted.

And hire 2 or 3 dev ops to keep it running ? That another 400 to 700k. It becomes your problem and not someone else’s. However, I don’t trust hosted LLMs for anything that needs to be private.

Where do I sign up to get 200k/yr to keep one rack running? Sounds like an incredibly chill job

What you get is not what you cost.

40% overhead is quite typical, so you'd be looking at $120k/year. In the USA I'd consider that a competitive salary for an admin capable or keeping a $6M rack of specialized hardware running 24/7.

Re: Kimi-K3 Technical Report [pdf]

#109
post #74

Earlier quoted context omitted.

> Kudos to Moonshot for making these weights available for download. Lets not fool ourselves and claim these are "open source" by any understanding of the concept though, there are usage restrictions (even if you download them) and also training data isn't clearly broken down either, nor it it actually using a FOSS license. They will probably never release the training data because that represents a large part of the…

Exactly. I wish people would just stop saying "open source" regarding models. Even "open weight" is disingenuous. "self-hostable" would be more honest.

It's a gradient almost, with some steps. So far, I think you could categorize every single released so far as one of:

- Proprietary - No access beyond remote endpoints

- Downloadable - You can run it, but there are restrictions and training data/code isn't public and/or under FOSS license, nor are the weights under a FOSS license

- Open weights - The weights are under a FOSS license and downloadable without restrictions, but not all of training code/data is public or under a FOSS license

- Open source - The model architecture, weights, training code and data are all under a FOSS license, and you can freely download and use them without any restrictions

Few models hit the mark for "open source model", many models people call "open weights" would go under "downloadable" here, which personally I think would be accurate.

Re: Kimi-K3 Technical Report [pdf]

#110
post #78

Earlier quoted context omitted.

only if you only get your news from main stream business press and Big Lab propaganda channels There's no chance K3 is a distill of Fable, it came out way too soon after the limited fable release to be feasbile. If you look at all of the top ML conferences, chinese labs contribute way more to advances in ML than "Open"AI and Anthropic: https://www.reddit.com/r/TheMachineGod/comments/1pi4q7f/pape... This K3 release ju…

Appreciate this point. Back in grad school, I published a peer reviewed paper with all the source code and datasets used. Got heckled at a conference talk by staff from a commercial lab. They shouted, we figured this out five years ago, lol. Also our approach is still better. But they don’t release their source code or publish much, so no one knows what this approach is or if it’s actually better. And years down the…

Oh this definitely happens all the time. I was an early employee at Clarifai, which won imagenet a year after alexnet and we were able to stay on the frontier for about 2 years before a bunch of open source models were matching our results. It was always some random PhD research project spinout or some random kid in Boston named Alec Radford.

We had a bunch of things that we never published that ended up being major research findings years later at top conferences.

Post reply on HN