Live data from Hacker News

Kimi K2 1T model runs on 2 512GB M3 Ultras

twitter.com

41–50 of 125 posts

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#41

Earlier quoted context omitted.

As long as you're willing to wait up to an hour for your GPU to get scheduled when you do want to use it.

I don’t understand what you’re saying. What’s preventing you from using eg OpenRouter to run a query against Kimi-K2 from whatever provider?

and you'll get a faster model this way

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#42

Earlier quoted context omitted.

> As a chatbot, it's the only one that seems to really relish calling you out on mistakes or nonsense, and it doesn't hesitate to be blunt with you. My experience is that Sonnet 4.5 does this a lot as well, but this is more often than not due to a lack of full context, eg accusing the user of not doing X or Y when it just wasn’t told that was already done, and proceeding to apologize. How is Kimi K2 in this regard? I…

> Isn’t “instruction following” the most important thing you’d want out of a model in general, No. And for the same reason that pure "instruction following" in humans is considered a form of protest/sabotage. https://en.wikipedia.org/wiki/Work-to-rule

It's still insanity to me that doing your job exactly as defined and not giving away extra work is considered a form of action.

Everyone should be working-to-rule all the time.

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#43

Kimi K2 is a really weird model, just in general. It's not nearly as smart as Opus 4.5 or 5.2-Pro or whatever, but it has a very distinct writing style and also a much more direct "interpersonal" style. As a writer of very-short-form stuff like emails, it's probably the best model available right now. As a chatbot, it's the only one that seems to really relish calling you out on mistakes or nonsense, and it doesn't h…

It's also the only model that consistently nails my favorite AI benchmark: https://clocks.brianmoore.com/

I use that one for image gen too. Ask for a picture of a grandfather clock at a specific time. Most are completely unable. Clocks are always 10:20 because that's the most photogenic time used in most stock photos.

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#45

I get tempted to buy a couple of these, but I just feel like the amortization doesn’t make sense yet. Surely in the next few years this will be orders of magnitude cheaper.

Before committing to purchasing two of these, you should look at the true speeds that few people post. Not just the "it works". We're at a point where we can run these very large models "at home", and it is great! But true usage is now with very large contexts, both in prompt processing, and token generations. Whatever speeds these models get at "0" context is very different than what they get at "useful" context, especially in coding and such.

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#46

Earlier quoted context omitted.

> Isn’t “instruction following” the most important thing you’d want out of a model in general, No. And for the same reason that pure "instruction following" in humans is considered a form of protest/sabotage. https://en.wikipedia.org/wiki/Work-to-rule

I don’t understand the point you’re trying to make. LLMs are not humans. From my perspective, the whole problem with LLMs (at least for writing code) is that it shouldn’t assume anything, follow the instructions faithfully, and ask the user for clarification if there is ambiguity in the request. I find it extremely annoying when the model pushes back / disagrees, instead of asking for clarification. For this reason,…

> and ask the user for clarification if there is ambiguity in the request.

You'd just be endlessly talking to the chatbots. Humans are really bad at expressing ourselves precisely, which is why we have formal languages that preclude ambiguity.

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#47

Earlier quoted context omitted.

I can't help you then. You can find a close analogue in the OSS/CIA Simple Sabotage Field Manual. [1] For that reason, I don't trust Agents (human or ai, secret or overt :-P) who don't push back. [1] https://www.cia.gov/static/5c875f3ec660e092cf893f60b4a288df/... esp. Section 5(11)(b)(14): "Apply all regulations to the last letter." - [as a form of sabotage]

How is asking for clarification before pushing back a bad thing?

Sounds like we're not too far apart then!

Sometimes pushback is appropriate, sometimes clarification. The key thing is that one doesn't just blindly follow instructions; at least that's the thrust of it.

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#48

Earlier quoted context omitted.

As long as you're willing to wait up to an hour for your GPU to get scheduled when you do want to use it.

I don’t understand what you’re saying. What’s preventing you from using eg OpenRouter to run a query against Kimi-K2 from whatever provider?

Because you have Cloudflare (MITM 1), Openrouter (MITM 2) and finally the "AI" provider who can all read, store, analyze and resell your queries.

EDIT: Thanks for downvoting what is literally one of the most important reasons for people to use local models. Denying and censoring reality does not prevent the bubble from bursting.

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#49

Earlier quoted context omitted.

I don’t think it will ever make sense; you can buy so much cloud based usage for this type of price. From my perspective, the biggest problem is that I am just not going to be using it 24/7. Which means I’m not getting nearly as much value out of it as the cloud based vendors do from their hardware. Last but not least, if I want to run queries against open source models, I prefer to use a provider like Groq or Cerebr…

my issue is once you have it in your workflow I'd be pretty latency sensitive. imagine those record-it-all apps working well. eventually you'd become pretty reliant on it. I don't want to necessarily be at the whims of the cloud

Aren’t those “record it all” applications implemented as a RAG and injected into the context based on embedding similarity?

Obviously you’re not going to always inject everything into the context window.

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#50

Kimi K2 is a really weird model, just in general. It's not nearly as smart as Opus 4.5 or 5.2-Pro or whatever, but it has a very distinct writing style and also a much more direct "interpersonal" style. As a writer of very-short-form stuff like emails, it's probably the best model available right now. As a chatbot, it's the only one that seems to really relish calling you out on mistakes or nonsense, and it doesn't h…

> I get the feeling that it was trained very differently from the other models

It's actually based on a deepseek architecture just bigger size experts if I recall correctly.

Post reply on HN