Live data from Hacker News

Kimi-K3 Technical Report [pdf]

github.com

121–130 of 199 posts

Re: Kimi-K3 Technical Report [pdf]

#121
post #5

Also open sourced a bunch of infra to go with it. Anyone who claims open source and open weights models are "decel" needs to get their head checked https://github.com/MoonshotAI/MoonEP https://github.com/kvcache-ai/AgentEnv https://github.com/MoonshotAI/FlashKDA

This comment would be much better without the second line

It’s a different topic but it’s correct imo. Open sourcing things is the best way to accelerate development.

Put another way, if you want to slow things down, put it behind a paywall, tag ideas ans “intellectual property” (meaning you’re the only one who can use it) and get the lawyers involved (injecting our slow legal system).

None of the above is a judgement call on whether development should be accelerated.

Re: Kimi-K3 Technical Report [pdf]

#122
post #108

Earlier quoted context omitted.

Where do I sign up to get 200k/yr to keep one rack running? Sounds like an incredibly chill job

What you get is not what you cost. 40% overhead is quite typical, so you'd be looking at $120k/year. In the USA I'd consider that a competitive salary for an admin capable or keeping a $6M rack of specialized hardware running 24/7.

Yeah but you don't need two such people, or even one, dedicated to this single rack.

A company of the size that this is worthwhile for, probably has dedicated devops on staff already and can add this rack to the inventory with no additional staff.

Re: Kimi-K3 Technical Report [pdf]

#123
post #50

Earlier quoted context omitted.

This comment would be much better without the second line

I know it's a hard ask on this site, but I need you to start parsing content and not tone. It was a helpful bit of context, even if it was a bit vitriolic.

Some battles are simply not going to be won.

I, for example, dislike reading comments complaining the submission (or another comment) is LLM generated. Focus on the content, not the style.

I'm not going to have my way, and nor shall you.

Re: Kimi-K3 Technical Report [pdf]

#124
post #108

Earlier quoted context omitted.

What you get is not what you cost. 40% overhead is quite typical, so you'd be looking at $120k/year. In the USA I'd consider that a competitive salary for an admin capable or keeping a $6M rack of specialized hardware running 24/7.

Yeah but you don't need two such people, or even one, dedicated to this single rack. A company of the size that this is worthwhile for, probably has dedicated devops on staff already and can add this rack to the inventory with no additional staff.

Who is going to upgrade the models ?

Who is going to fix it when the api does something weird ?

Who is going to proactively make sure it’s not overheating?

Chat GPT has enterprise contracts for a reason.

Re: Kimi-K3 Technical Report [pdf]

#126
post #120
post #104

Earlier quoted context omitted.

From the other dev-op work you’re doing.

Understood but sharing your existing devops resources with this will soon become a bottleneck especially when any major downtime will keep several engineers (and long-running agents) blocked from any meaningful work until availability improves.

Not my experience, from an SMB that maintains its own hardware and services. You have a certain contingent of competent engineers who distribute their work across projects, and it generally works out fine. Or course you plan with some redundancy and fall-back plans in your systems.

Re: Kimi-K3 Technical Report [pdf]

#127
post #9

Earlier quoted context omitted.

This comment would be much better without the second line

I'm not sure I understand the case for open-source models being decelerationist, is this it? Decel: - Potentially reduces investor appetite for funding big labs. - More risk of powerful AI getting in bad hands -> more regulation. Accel: - More competition so big labs can't rest on laurels. - More research in open, so all labs can accrete advancements faster. I feel like open-source = acceleration has a much more clea…

the argument is that we should all fold and let sam altman burn trillions of dollars on naive scaling and pay monopoly prices for their closed APIs until the models are good enough to be closed off for "safety" reasons so that they can take an even larger cut by competing directly with us

Re: Kimi-K3 Technical Report [pdf]

#128

Earlier quoted context omitted.

Yeah but you don't need two such people, or even one, dedicated to this single rack. A company of the size that this is worthwhile for, probably has dedicated devops on staff already and can add this rack to the inventory with no additional staff.

Who is going to upgrade the models ? Who is going to fix it when the api does something weird ? Who is going to proactively make sure it’s not overheating? Chat GPT has enterprise contracts for a reason.

[deleted]

Re: Kimi-K3 Technical Report [pdf]

#129

Back of the envelope calculation (could be off, correct me if I am) If you are a large enough company that spends million+ on inference a month, it makes sense to buy a GB300 rack ($6M on top range from what I could find) which has 20.7 TB. Since the model is mixed trained (MXFP4), you would need less than 10% of the rack's memory to serve the full model. Aggregate HBM bandwidth: 576 TB/s. You can run over 6000 paral…

You also need a place to put it. With liquid cooling and extremely dense power capability. Your typical colo or closet server room isn't going to cut it.

Re: Kimi-K3 Technical Report [pdf]

#130

License: https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE > If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or…

What is it a license for, though? Copyright? Are model weights even copyrightable?

These are the right questions. IANAL, but I am of the opinion that model weights are not copyrightable. See my sibling comment: https://news.ycombinator.com/item?id=49073127
Post reply on HN