Live data from Hacker News

Kimi K2 1T model runs on 2 512GB M3 Ultras

twitter.com

71–80 of 125 posts

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#71
post #65

Earlier quoted context omitted.

Careful with that benchmark. It's LLMs grading other LLMs.

Well if lmsys showed anything, it's that human judges are measurably worse. Then you have your run of the mill multiple choice tests that grade models on unrealistic single token outputs. What does that leave us with?

Seems like a foreshock of AGI if the average human is no longer good enough to give feedback directly and the nets instead have to do recursive self improvement themselves.

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#72

I get tempted to buy a couple of these, but I just feel like the amortization doesn’t make sense yet. Surely in the next few years this will be orders of magnitude cheaper.

Before committing to purchasing two of these, you should look at the true speeds that few people post. Not just the "it works". We're at a point where we can run these very large models "at home", and it is great! But true usage is now with very large contexts, both in prompt processing, and token generations. Whatever speeds these models get at "0" context is very different than what they get at "useful" context, es…

Are there benchmarks that effectively measure this? This is essential information when speccing out an inference system/model size/quantization type.

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#73
post #52

Kimi K2 is a really weird model, just in general. It's not nearly as smart as Opus 4.5 or 5.2-Pro or whatever, but it has a very distinct writing style and also a much more direct "interpersonal" style. As a writer of very-short-form stuff like emails, it's probably the best model available right now. As a chatbot, it's the only one that seems to really relish calling you out on mistakes or nonsense, and it doesn't h…

In their AMA moonshot said it was mainly finetuning

OpenAI and the other big players clearly RLHF with different users in mind than professionals. They’re optimizing for sycophancy and general pleasantness. It’s beautiful to finally see a big model that hasn’t been warped in this way. I want a model that is borderline rude in its responses. Concise, strict, and as distrustful of me as I am of it.

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#74

Kimi K2 is a really weird model, just in general. It's not nearly as smart as Opus 4.5 or 5.2-Pro or whatever, but it has a very distinct writing style and also a much more direct "interpersonal" style. As a writer of very-short-form stuff like emails, it's probably the best model available right now. As a chatbot, it's the only one that seems to really relish calling you out on mistakes or nonsense, and it doesn't h…

It's a lot stronger for geospatial intelligence tasks than any other model in my experience. Shame it's so slow in terms of tps

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#76
post #50

Kimi K2 is a really weird model, just in general. It's not nearly as smart as Opus 4.5 or 5.2-Pro or whatever, but it has a very distinct writing style and also a much more direct "interpersonal" style. As a writer of very-short-form stuff like emails, it's probably the best model available right now. As a chatbot, it's the only one that seems to really relish calling you out on mistakes or nonsense, and it doesn't h…

> I get the feeling that it was trained very differently from the other models It's actually based on a deepseek architecture just bigger size experts if I recall correctly.

It was notably trained with Muon optimizer for what it's worth, but I don't know how much can be attributed to that alone

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#77

Earlier quoted context omitted.

I think you’re missing the whole point, which is not using cloud compute.

Because of privacy reasons? Yeah I’m not going to spend a small fortune for that to be able to use these types of models.

There are plenty of examples and reasons to do so besides privacy- because one can, because it’s cool, for research, for fine tuning, etc. I never mentioned privacy. Your use case is not everyone’s.

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#78

Kimi K2 is a very impressive model! It's particularly un-obsequious, which makes it useful for actually checking your reasoning on things. Some especially older ChatGPT models will tell you that everything you say is fantastic and great. Kimi -on the other hand- doesn't mind taking a detour to question your intelligence and likely your entire ancestry if you ask it to be brutal.

I made the mistake of turning off nsfw mode while in a buddy's Tesla and then Grok misheard something else I said as "I like lesbians", and it just went off on me. It was pretty hilarious. That model is definitely not obsequious either.

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#79
post #23

Earlier quoted context omitted.

Autonomy generally, not just privacy. You never know what the future will bring, AI will be enshittified and so will hubs like huggingface. It’s useful to have an off grid solution that isn’t subject to VCs wanting to see their capital returned.

> You never know what the future will bring, AI will be enshittified and so will hubs like huggingface. If anyone wants to bet that future cloud hosted AI models will get worse than they are now, I will take the opposite side of that bet. > It’s useful to have an off grid solution that isn’t subject to VCs wanting to see their capital returned. You can pay cloud providers for access to the same models that you can ru…

If anyone wants to bet that future cloud hosted AI models will get worse than they are now, I will take the opposite side of that bet.

OK. How do we set up this wager?

I'm not knowledgeable about online gambling or prediction markets, but further enshittification seems like the world's safest bet.

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#80
post #38

Earlier quoted context omitted.

How do you feel K2 Thinking compares to Opus 4.5 and 5.2-Pro?

? The user directly addresses this.

K2 and K2T are drastically different models released a significant amount of time apart, with wildly different capabilities and post training. K2T is much closer in capability to 4.5 Sonnet from what I've heard.
Post reply on HN