Live data from Hacker News

Kimi-K3 Technical Report [pdf]

github.com

81–90 of 199 posts

Re: Kimi-K3 Technical Report [pdf]

#81

Back of the envelope calculation (could be off, correct me if I am) If you are a large enough company that spends million+ on inference a month, it makes sense to buy a GB300 rack ($6M on top range from what I could find) which has 20.7 TB. Since the model is mixed trained (MXFP4), you would need less than 10% of the rack's memory to serve the full model. Aggregate HBM bandwidth: 576 TB/s. You can run over 6000 paral…

And hire 2 or 3 dev ops to keep it running ? That another 400 to 700k. It becomes your problem and not someone else’s. However, I don’t trust hosted LLMs for anything that needs to be private.

There will be cloud/SaaS vendors who have lower cost of labor/capital due to automation and financing terms.

Having these models in the open caps the inference margin.

Re: Kimi-K3 Technical Report [pdf]

#82

Earlier quoted context omitted.

And hire 2 or 3 dev ops to keep it running ? That another 400 to 700k. It becomes your problem and not someone else’s. However, I don’t trust hosted LLMs for anything that needs to be private.

Where do I sign up to get 200k/yr to keep one rack running? Sounds like an incredibly chill job

Apparently it’s going to take the 3 of us to do this, mate. Going to get so much reading done.

Re: Kimi-K3 Technical Report [pdf]

#83
post #9

Earlier quoted context omitted.

I'm not sure I understand the case for open-source models being decelerationist, is this it? Decel: - Potentially reduces investor appetite for funding big labs. - More risk of powerful AI getting in bad hands -> more regulation. Accel: - More competition so big labs can't rest on laurels. - More research in open, so all labs can accrete advancements faster. I feel like open-source = acceleration has a much more clea…

If your worldview is “most of the progress is made by closed labs, then open labs fast-follow” (which isn’t implausible given the documented distillation of Fable), and further that open labs cannot make make meaningful progress vs the closed labs except by fast-following and that they won’t pick up the ability to make progress after the closed labs are gone, then driving closed labs out of business slows down overal…

I think it's pretty hard to hold that worldview: Anthropic couldn't ship a reasoning model until they copied DeepSeek R1's homework, and they've all copied DS-style super-sparse MoEs at this point too.

Re: Kimi-K3 Technical Report [pdf]

#84

Earlier quoted context omitted.

And hire 2 or 3 dev ops to keep it running ? That another 400 to 700k. It becomes your problem and not someone else’s. However, I don’t trust hosted LLMs for anything that needs to be private.

> However, I don’t trust hosted LLMs for anything that needs to be private. Why not? Do you trust AWS with things that need to be private?

More than I trust frontier labs. AWS doesn't need to recoup 9 digits USD of capex

Re: Kimi-K3 Technical Report [pdf]

#85
post #61
post #33

Earlier quoted context omitted.

Kimi 2.6 had the same janky license: https://huggingface.co/moonshotai/Kimi-K2.6/blob/main/LICENS... - looks like they've been doing that at least as far back as K2. The models they released in 2025 - https://huggingface.co/moonshotai/models - were clean MIT. They started doing the "modified MIT" thing in January 2026 with moonshotai/Kimi-K2-Thinking

In what way was the license "janky"? Moonshot isn't a trillion dollar US behemoth. If these terms allow it to release near frontier models with open weights, so startups and anyone with enough hardware can use them, while charging only companies with more than $20 million in revenue or 100 million users, that seems like a reasonable trade off.

I use the term "janky" for any time someone releases something under a supposedly open source license that doesn't comply with the OSI definition.

"Modified MIT" is the perfect example of that.

I'm not saying it's unreasonable, or that you can't release under such a license - it's your software, use whatever license you like!

I'll call it "janky" when you do.

Re: Kimi-K3 Technical Report [pdf]

#86

Back of the envelope calculation (could be off, correct me if I am) If you are a large enough company that spends million+ on inference a month, it makes sense to buy a GB300 rack ($6M on top range from what I could find) which has 20.7 TB. Since the model is mixed trained (MXFP4), you would need less than 10% of the rack's memory to serve the full model. Aggregate HBM bandwidth: 576 TB/s. You can run over 6000 paral…

I think the licensing that would likely apply to a company that's able to afford ~$6M rack and the associated infrastructure muddies this somewhat

I think internal use is allowed at any scale in the license?

> 4. The requirements set forth in Sections 2 and 3 do not apply to: (a) internal use of the Software, defined as any use that does not make the Software, its outputs, or its underlying capabilities available to third parties; [...]

Re: Kimi-K3 Technical Report [pdf]

#87
post #50

Earlier quoted context omitted.

This comment would be much better without the second line

I know it's a hard ask on this site, but I need you to start parsing content and not tone. It was a helpful bit of context, even if it was a bit vitriolic.

The content was 'you need to get your head checked'. That isn't tone, that directly implying that if you hold that position, there is something wrong with you. It's rude and unnecessary.

Re: Kimi-K3 Technical Report [pdf]

#88

Back of the envelope calculation (could be off, correct me if I am) If you are a large enough company that spends million+ on inference a month, it makes sense to buy a GB300 rack ($6M on top range from what I could find) which has 20.7 TB. Since the model is mixed trained (MXFP4), you would need less than 10% of the rack's memory to serve the full model. Aggregate HBM bandwidth: 576 TB/s. You can run over 6000 paral…

And hire 2 or 3 dev ops to keep it running ? That another 400 to 700k. It becomes your problem and not someone else’s. However, I don’t trust hosted LLMs for anything that needs to be private.

Just like that new jobs created by AI! Localized model maintainer/technician.

Re: Kimi-K3 Technical Report [pdf]

#89
post #50

Earlier quoted context omitted.

I know it's a hard ask on this site, but I need you to start parsing content and not tone. It was a helpful bit of context, even if it was a bit vitriolic.

Why should I waste time parsing content and not tone? Why can't the commenter just avoid the tone? It even saves time since you can write less!

Because discourse around tone isn't productive. You could wipe this whole comment thread, starting with the parent of mine, and lose exactly zero information.

Re: Kimi-K3 Technical Report [pdf]

#90

Back of the envelope calculation (could be off, correct me if I am) If you are a large enough company that spends million+ on inference a month, it makes sense to buy a GB300 rack ($6M on top range from what I could find) which has 20.7 TB. Since the model is mixed trained (MXFP4), you would need less than 10% of the rack's memory to serve the full model. Aggregate HBM bandwidth: 576 TB/s. You can run over 6000 paral…

And hire 2 or 3 dev ops to keep it running ? That another 400 to 700k. It becomes your problem and not someone else’s. However, I don’t trust hosted LLMs for anything that needs to be private.

You’ll slap some training on existing technologists/infra/sysadmin folks and perhaps have a support contract for the edge cases (hardware troubleshooting and advanced replacement).

(managed an entire data center building with thousands of servers a lifetime ago with ~2-3 other people, it’s only gotten easier over the last two decades imho)

Post reply on HN