Live data from Hacker News

Kimi-K3 Technical Report [pdf]

github.com

171–180 of 199 posts

Re: Kimi-K3 Technical Report [pdf]

#171
post #50

Earlier quoted context omitted.

I know it's a hard ask on this site, but I need you to start parsing content and not tone. It was a helpful bit of context, even if it was a bit vitriolic.

Some battles are simply not going to be won. I, for example, dislike reading comments complaining the submission (or another comment) is LLM generated. Focus on the content, not the style. I'm not going to have my way, and nor shall you.

Oh, I understand. It's the nature of the voting system of comment feedback. Reddit behaves the same way. Arguments become competitions to see who can inject enough vitriol while still maintaining a placid genteel demeanor. The one who "wins" (convinces the peanut gallery to upvote them) is the one who can avoid looking like they got mad.

It's why people can advocate for ethnic cleansing here, and that's fine as long as they word it correctly, but if someone calls them an asshole about it, they're flagged.

Really common in rationalist circles from my experience as well because they believe that true statements aren't always normative, and that, since their arguments are true because they're rational, their statements aren't necessarily normative. Begging the question, of course, but I see that in situations like Scott Alexander's defense of "human biodiversity" theories (the whole HBD moniker is itself an example of everything I'm talking about condensed into two words).

Re: Kimi-K3 Technical Report [pdf]

#172
post #171

Earlier quoted context omitted.

Some battles are simply not going to be won. I, for example, dislike reading comments complaining the submission (or another comment) is LLM generated. Focus on the content, not the style. I'm not going to have my way, and nor shall you.

Oh, I understand. It's the nature of the voting system of comment feedback. Reddit behaves the same way. Arguments become competitions to see who can inject enough vitriol while still maintaining a placid genteel demeanor. The one who "wins" (convinces the peanut gallery to upvote them) is the one who can avoid looking like they got mad. It's why people can advocate for ethnic cleansing here, and that's fine as long…

> Oh, I understand. It's the nature of the voting system of comment feedback. Reddit behaves the same way. Arguments become competitions to see who can inject enough vitriol while still maintaining a placid genteel demeanor. The one who "wins" (convinces the peanut gallery to upvote them) is the one who can avoid looking like they got mad.

Depends on which subreddit and which flavor of groupthink. The behavior you say is upvoted is one I often see downvoted to oblivion on Reddit.

> but if someone calls them an asshole about it, they're flagged.

That's because name calling is against HN guidelines.

Just an aside since it's not clear: I think the asshole is the one using labels like "decel" (or even "MAGA" unless the person self describes). It's irrelevant if I agree with the rest of the comment.

It's OK to call people out for name calling while still agreeing with the rest. It's problematic to require one acknowledge the quality of the rest of the comment when calling them out on their name calling.

Put in a less twisted manner: If it's OK to address his comment sans the "decel", it should also be OK to address his use of "decel" without discussing the rest of the comment.

Re: Kimi-K3 Technical Report [pdf]

#173

Earlier quoted context omitted.

Demonstrating a company is using this model in court seems non-trivial, no? Or am I interpreting this post incorrectly.

Not at all. It is easy to bake in a specific response to a given prompt in the model.

It is trivial to contrive a distinct model that behaves the same, no?

Regardless, most uses of a model won't provide direct interaction. This entire line of thought seems starved of use. Unless we're talking about LLM providers themselves... but that is a separate topic entirely

Re: Kimi-K3 Technical Report [pdf]

#175

Back of the envelope calculation (could be off, correct me if I am) If you are a large enough company that spends million+ on inference a month, it makes sense to buy a GB300 rack ($6M on top range from what I could find) which has 20.7 TB. Since the model is mixed trained (MXFP4), you would need less than 10% of the rack's memory to serve the full model. Aggregate HBM bandwidth: 576 TB/s. You can run over 6000 paral…

I think the licensing that would likely apply to a company that's able to afford ~$6M rack and the associated infrastructure muddies this somewhat

Its an open question if model weights are even copyrightable. So the license may not mean anything.

Re: Kimi-K3 Technical Report [pdf]

#176

Earlier quoted context omitted.

Not at all. It is easy to bake in a specific response to a given prompt in the model.

It is trivial to contrive a distinct model that behaves the same, no? Regardless, most uses of a model won't provide direct interaction. This entire line of thought seems starved of use. Unless we're talking about LLM providers themselves... but that is a separate topic entirely

no, because you don't know what the "secret prompt" is. there could be hundreds such prompts

Re: Kimi-K3 Technical Report [pdf]

#177

Earlier quoted context omitted.

> And hire 2 or 3 dev ops to keep it running Not a devops but I'd say one full time is already too many.

Yes but zero is not enough and where do you get a fraction of a competent dev op from?

colocate and split the cost with others.

Re: Kimi-K3 Technical Report [pdf]

#178

Earlier quoted context omitted.

It is trivial to contrive a distinct model that behaves the same, no? Regardless, most uses of a model won't provide direct interaction. This entire line of thought seems starved of use. Unless we're talking about LLM providers themselves... but that is a separate topic entirely

no, because you don't know what the "secret prompt" is. there could be hundreds such prompts

Ok, I'll take this at face value... why would someone share this seemingly secret prompt with the prosecution?

Re: Kimi-K3 Technical Report [pdf]

#179

Earlier quoted context omitted.

The same person managing the company's email accounts and whatnot. I didn't say it's fire-and-forget. I'm saying all that is maybe a day of work every 3 months.

It doesn’t really work like that. The companies which have the will and the budget to host their own LLMs typically require a ton of other, much smaller, models as well. They have internal security requirements, guardrails, audit, critical workflows start depending on your onprem setup, downtime is now something that’s not even allowed. There’s going to be a zoo of tooling, lots of bespoke work with internal clients…

That's some of those companies not all of them.

I agree that what you've described is probably worth a dedicated person. The earlier comment describing one rack running one model, is not.

Re: Kimi-K3 Technical Report [pdf]

#180
post #85

Earlier quoted context omitted.

I use the term "janky" for any time someone releases something under a supposedly open source license that doesn't comply with the OSI definition. "Modified MIT" is the perfect example of that. I'm not saying it's unreasonable, or that you can't release under such a license - it's your software, use whatever license you like! I'll call it "janky" when you do.

So it's more like your personal definition, not the normal meaning of the word. Ok, because "janky" is informal slang also meaning poor quality, unreliable etc. I understand that you're using janky as shorthand for not OSI compliant. I just think that can blur two separate issues: whether a license qualifies as open source under the OSI definition, and whether the license itself is unreasonable/badly designed.

[deleted]
Post reply on HN