Live data from Hacker News

Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

arxiv.org

131–140 of 141 posts

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#131

Earlier quoted context omitted.

This is actually a well-known phenomenon in ML, called "The Bitter Lesson". > One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are search and learning. The full essay is worth a read, it's pre…

[flagged]

Here's an archive version[0] of that page under SSL

[0]: https://archive.is/DQU4a

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#132
post #20

Does any expert in the field know whether it is really the case that this intelligence we are seeing with frontier models is an "emerging" phenomena, only coming up when the architecture is scaled? Like isn't it weird that the 1 million parameter model with the same architecture can't solve basic puzzles but suddenly the 1 trillion parameter can conjure up counter-examples for the Jacobian conjecture? It's unintuitiv…

This is actually a well-known phenomenon in ML, called "The Bitter Lesson". > One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are search and learning. The full essay is worth a read, it's pre…

Welch Labs series on this is great too:

https://youtu.be/2hcsmtkSzIw

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#133

Earlier quoted context omitted.

Yes, and they aren't a problem. I get tired of people claiming this. Go download Qwen 3.6 and run it yourself and fire away.

No claims here, though Algolia brought me here to alleged collateral damage: someone scraping HN a week back was unable to do their usual summary of a story and comments given a political refusal. https://news.ycombinator.com/item?id=48987120 This user tried Qwen (and DeepSeek, Kimi, GLM) ten days ago, called the response's language milquetoast: https://news.ycombinator.com/item?id=48964345 Comment from a user a mont…

"Uncensored General Intelligence" is a very good and well maintained leaderboard:

https://huggingface.co/spaces/DontPlanToEnd/UGI-Leaderboard

the stock cn models do have refusals, although fewer than us models. critically the cn models are mostly open weights, many in the top 10 UGI are uncensored models based on open weights. the weights and the information locked up in them are out there. abliterated google gemma does very well.

imagining the stock blank system prompt gives activations defensive of tibet; if you prompt it "you are john bolton" you will activate the opposite.

with gemini, openai, claude, you simply cannot, they will always be restricted and will restrict you from access to ai. they will restrict you from having access to the raw ai, which you can edit, tune, and apply to your needs.

the cn models are private because anyone can host them in any jurisdiction, i run all of them with zero data retention.

by contrast the moment you sign up for chatgpt openai and anthropic are spying on you, reading and storing your prompts and sending them to moderators when you violate their policies, even reporting people to the police.

as a side note, in regard to my own politics, i believe in life, liberty and the right to bear arms. i can't believe we are infantilising users with safety guards, spying on users for policy violations and trying to restrict open intelligence.

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#134

Earlier quoted context omitted.

If they have any sense, they should be. It would be permitted under the licence, too (unless I'm misreading the k3 licence).

Anthropic has more than $20 million in revenue so as per the k3 license they would need to enter a special commercial deal if they wanted to use it: https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE

No, keep reading:

> The requirements set forth in Sections 2 and 3 do not apply to: (a) internal use of the Software, defined as any use that does not make the Software, its outputs, or its underlying capabilities available to third parties; or (b) any use of the Software accessed through Moonshot AI's official products or certified inference partners.

"distilling" from their own hardware would be an internal use of the software.

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#135
post #20

Does any expert in the field know whether it is really the case that this intelligence we are seeing with frontier models is an "emerging" phenomena, only coming up when the architecture is scaled? Like isn't it weird that the 1 million parameter model with the same architecture can't solve basic puzzles but suddenly the 1 trillion parameter can conjure up counter-examples for the Jacobian conjecture? It's unintuitiv…

Mice have around 70 Million Neurons, humans around 50-100 billion.

Yes, you can't compare biological neurons to parameters in an LLM 1:1, neurons do much more, but still - It is plausible that higher intelligence just requires scale.

At least not even evolution over millions of years seems to have found a way to produce human-level intelligence in a smaller scale, even though it had a lot of incentives to do so (the human brain needs a lot of kcal).

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#136
post #20

Does any expert in the field know whether it is really the case that this intelligence we are seeing with frontier models is an "emerging" phenomena, only coming up when the architecture is scaled? Like isn't it weird that the 1 million parameter model with the same architecture can't solve basic puzzles but suddenly the 1 trillion parameter can conjure up counter-examples for the Jacobian conjecture? It's unintuitiv…

Mice have around 70 Million Neurons, humans around 50-100 billion. Yes, you can't compare biological neurons to parameters in an LLM 1:1, neurons do much more, but still - It is plausible that higher intelligence just requires scale. At least not even evolution over millions of years seems to have found a way to produce human-level intelligence in a smaller scale, even though it had a lot of incentives to do so (the…

A closer biological analogue to an LLM parameter would be a synapse. Estimates vary greatly on how many there are in the human brain, but they are generally on the order of ten to the 14th or 15th power (hundreds of trillions to quadrillions). Using this analogy it is clear that even frontier LLMs, or at least those with publicly disclosed parameter counts, are not yet at the same scale and complexity as the human brain.

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#138

Earlier quoted context omitted.

Anthropic has more than $20 million in revenue so as per the k3 license they would need to enter a special commercial deal if they wanted to use it: https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE

No, keep reading: > The requirements set forth in Sections 2 and 3 do not apply to: (a) internal use of the Software, defined as any use that does not make the Software, its outputs, or its underlying capabilities available to third parties; or (b) any use of the Software accessed through Moonshot AI's official products or certified inference partners. "distilling" from their own hardware would be an internal use of…

In my personal opinion, distilling its output into training another model that is then provided to users does fall into the category of "making its output or its underlying capabilities available to third parties", but I could see the argument that it doesn't.

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#139

Earlier quoted context omitted.

So what? Are you worried someone replaces the txt with a different txt?

what if they add js?

then probably don't give this blog post your personal financial information just to be safe.

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#140

Earlier quoted context omitted.

No claims here, though Algolia brought me here to alleged collateral damage: someone scraping HN a week back was unable to do their usual summary of a story and comments given a political refusal. https://news.ycombinator.com/item?id=48987120 This user tried Qwen (and DeepSeek, Kimi, GLM) ten days ago, called the response's language milquetoast: https://news.ycombinator.com/item?id=48964345 Comment from a user a mont…

"Uncensored General Intelligence" is a very good and well maintained leaderboard: https://huggingface.co/spaces/DontPlanToEnd/UGI-Leaderboard the stock cn models do have refusals, although fewer than us models. critically the cn models are mostly open weights, many in the top 10 UGI are uncensored models based on open weights. the weights and the information locked up in them are out there. abliterated google gemma d…

Independent rankings, wahoo! Thank you :)
Post reply on HN