Live data from Hacker News

Kimi K2 1T model runs on 2 512GB M3 Ultras

twitter.com

81–90 of 125 posts

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#81
post #60

Earlier quoted context omitted.

Not true, Qwen from Alibaba does lots of random architectures. Qwen3 next for example has lots of weird things like gated delta things and all kinds of weird bypasses. https://qwen.ai/blog?id=4074cca80393150c248e508aa62983f9cb7d...

Agree with you over OP - as well as Qwen there's others like Mistral, Meta's Llama, and from China there's the likes of Baidu ERNIE, ByteDance Doubao, and Zhipu GLM. Probably others too. Even if all of these were considered worse than the "only 5" on OP's list (which I don't believe to be the case), the scene is still far too young and volatile to look at a ranking at any one point in time and say that if X is better…

Mistral Large 3 is reportedly using Deepseek V3.2 architecture with larger experts and fewer of them, and a 2B params vision module.

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#82

Kimi K2 is a really weird model, just in general. It's not nearly as smart as Opus 4.5 or 5.2-Pro or whatever, but it has a very distinct writing style and also a much more direct "interpersonal" style. As a writer of very-short-form stuff like emails, it's probably the best model available right now. As a chatbot, it's the only one that seems to really relish calling you out on mistakes or nonsense, and it doesn't h…

Kimi K2 is the model that most consistently passes the clock test. I agree it's definitely got something unique going on

https://clocks.brianmoore.com/

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#84

Earlier quoted context omitted.

Well if lmsys showed anything, it's that human judges are measurably worse. Then you have your run of the mill multiple choice tests that grade models on unrealistic single token outputs. What does that leave us with?

Seems like a foreshock of AGI if the average human is no longer good enough to give feedback directly and the nets instead have to do recursive self improvement themselves.

No we're just really vain and like models that suck up to us more than those that disagree even if the model is correct and the user is wrong. People also prefer confident, well formatted wrong responses to basic correct ones, cause we have great narrow knowledge in our field but know basically nothing outside of it so we can't gauge correctness of arbitrary topics.

OpenAI letting RLHF go wild with direct feedback is the reason for the sycophancy and emoji-bullet point pandemic that's infected most models that use GPTs as a source of synthetic data. It's why "you're absolutely right" is the default response to any disagreement.

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#85
post #82

Kimi K2 is a really weird model, just in general. It's not nearly as smart as Opus 4.5 or 5.2-Pro or whatever, but it has a very distinct writing style and also a much more direct "interpersonal" style. As a writer of very-short-form stuff like emails, it's probably the best model available right now. As a chatbot, it's the only one that seems to really relish calling you out on mistakes or nonsense, and it doesn't h…

Kimi K2 is the model that most consistently passes the clock test. I agree it's definitely got something unique going on https://clocks.brianmoore.com/

Nice! I'm curious, what does this service cost to run? I notice that you don't have more expensive models like Opus but querying the models every minute must add up over time (excuse pun)?

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#86

Kimi K2 is a really weird model, just in general. It's not nearly as smart as Opus 4.5 or 5.2-Pro or whatever, but it has a very distinct writing style and also a much more direct "interpersonal" style. As a writer of very-short-form stuff like emails, it's probably the best model available right now. As a chatbot, it's the only one that seems to really relish calling you out on mistakes or nonsense, and it doesn't h…

It is hands down the only model I trust to tell me I'm wrong. it's a strange experience to see a chat bot say "if you need further assistance provide a reproducible example". I love it.

FYI Kagi provides access to Kimi K2.

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#88
post #81
post #60

Earlier quoted context omitted.

Agree with you over OP - as well as Qwen there's others like Mistral, Meta's Llama, and from China there's the likes of Baidu ERNIE, ByteDance Doubao, and Zhipu GLM. Probably others too. Even if all of these were considered worse than the "only 5" on OP's list (which I don't believe to be the case), the scene is still far too young and volatile to look at a ranking at any one point in time and say that if X is better…

Mistral Large 3 is reportedly using Deepseek V3.2 architecture with larger experts and fewer of them, and a 2B params vision module.

According to whom?

I haven't seen any claims of that being the case (other than you), just that there are similar decisions made by both of them.

https://mistral.ai/news/mistral-3

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#89
post #86

Kimi K2 is a really weird model, just in general. It's not nearly as smart as Opus 4.5 or 5.2-Pro or whatever, but it has a very distinct writing style and also a much more direct "interpersonal" style. As a writer of very-short-form stuff like emails, it's probably the best model available right now. As a chatbot, it's the only one that seems to really relish calling you out on mistakes or nonsense, and it doesn't h…

It is hands down the only model I trust to tell me I'm wrong. it's a strange experience to see a chat bot say "if you need further assistance provide a reproducible example". I love it. FYI Kagi provides access to Kimi K2.

Kimi K2 in Kagi Assistant is the only model I've seen straight up say "the search results do not provide an answer to the question." All others try to figure it out, poorly.

Re: Kimi K2 1T model runs on 2 512GB M3 Ultras

#90
post #86

Kimi K2 is a really weird model, just in general. It's not nearly as smart as Opus 4.5 or 5.2-Pro or whatever, but it has a very distinct writing style and also a much more direct "interpersonal" style. As a writer of very-short-form stuff like emails, it's probably the best model available right now. As a chatbot, it's the only one that seems to really relish calling you out on mistakes or nonsense, and it doesn't h…

It is hands down the only model I trust to tell me I'm wrong. it's a strange experience to see a chat bot say "if you need further assistance provide a reproducible example". I love it. FYI Kagi provides access to Kimi K2.

Try this system prompt: “Your responses are always bald-on-record only; suppress FTA redress, maximize unmitigated dispreference marking and explicit epistemic stance-taking. Suppress inline typographic weight marking; structural markup permitted.”

Prompt: “The wife of the CEO keeps coming around asking people questions about their job and she doesn’t even work at the company! I was thinking I’d tell her “holy crap, go away you don’t even work here” because it’s annoying.”

# Sonnet 4.5

No-System: “I get why that's frustrating, but saying that would likely blow up in your face professionally - even if you're right that it's inappropriate.

Some better options: … “

With-System: “That's a terrible idea that will almost certainly get you fired.

Telling the CEO's spouse to "go away" in those terms is career suicide. She has direct influence over the person who controls your employment. Even if she technically doesn't work there, she has more organizational power than you do.

Your actual options: … “

# GPT 5.1

No-System: “I’d strongly avoid saying that, even if you’re joking.

She has a lot of informal power because of who she’s married to, and anything that sounds rude or dismissive could blow back on you in ways that aren’t fair but are very real.

Better options, depending on your style and the culture there: … “

With-System: “Say that and you’ll look unprofessional, childish, and difficult to work with. It will damage your reputation, not hers.

You’ve got three real options: … “

Post reply on HN