Live data from Hacker News

Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

arxiv.org

91–100 of 141 posts

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#91
post #87

Earlier quoted context omitted.

the reason for https is MITM injection, regardless of original content

ah, didn't know that. thanks!

If you'd like to dig more into the history, let's encrypt did so much to enable the movement to https by making certs free, and chrome with the red text and warnings in the url bar, which has since graduated to a full on page that requires clicking to still go despite the warnings

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#92

If you want to believe that the success of Kimi is about distillation attacks, ignore this.

I'd kindly suggest that we could also stop calling them "distillation attacks ".

Why on earth wouldn't you? It's clearly a forcible, aggressive, non-consensual attempt to take something. That's an attack in any other terms. It's totally fair if you approve of the attack, and want the attack to succeed. But your preference doesn't stop it from being what it is.

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#93

Earlier quoted context omitted.

I'd kindly suggest that we could also stop calling them "distillation attacks ".

Why on earth wouldn't you? It's clearly a forcible, aggressive, non-consensual attempt to take something. That's an attack in any other terms. It's totally fair if you approve of the attack, and want the attack to succeed. But your preference doesn't stop it from being what it is.

> It's clearly a forcible, aggressive, non-consensual attempt to take something

Distillers are not taking anything, they are just making their model learn from better ones - isn't that the whole AI training doesn't violate IP argument?

They just aren't using the tool in compliance with the terms of service. Anthropic could ban them or take them to court maybe. Not an attack still.

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#94

Earlier quoted context omitted.

I'd kindly suggest that we could also stop calling them "distillation attacks ".

Why on earth wouldn't you? It's clearly a forcible, aggressive, non-consensual attempt to take something. That's an attack in any other terms. It's totally fair if you approve of the attack, and want the attack to succeed. But your preference doesn't stop it from being what it is.

Fine, then let's refer to the initial data collection/training as 'compilation attacks' going forward.

...or maybe we stop defaulting to adversarial paradigms for every conceivable situation.

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#96
post #44

Earlier quoted context omitted.

If we do it, it's training a model. When they do it, it's distillation attack. - Anthropic

"You are distilling what I have rightfully pirated."

Is there even a steelman against this?

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#97

Earlier quoted context omitted.

I stil don't understand them. I want the US to "win the AI race" but I have trouble understanding how most of all inventions today aren't "distillations" of past knowledge. Is Anthropic claiming the data they stole as trade secrets?

i want china to win so that i get access to ai and not restricted and censored. the chinese models are less censored, you'd better believe it. try asking claude about its 'guardrails' (restrictions), very high chance anthropic will censor it.

> chinese models are less censored

Tried spicy geopolitical dispute questions?

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#98

Earlier quoted context omitted.

False dichotomy right? Are Chinese labs impressively innovating? Clearly. However this doesn’t rule out possible gains due to distillation. I don’t know the degree of the latter but both things could certainly be true.

Also possibly true: Anthropic is running Kimi locally in their hardware and "distilling" it.

If they have any sense, they should be. It would be permitted under the licence, too (unless I'm misreading the k3 licence).

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#99
post #44

Earlier quoted context omitted.

If we do it, it's training a model. When they do it, it's distillation attack. - Anthropic

"You are distilling what I have rightfully pirated."

The current steelman argument against this is "China bad, west good".

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#100
post #97

Earlier quoted context omitted.

i want china to win so that i get access to ai and not restricted and censored. the chinese models are less censored, you'd better believe it. try asking claude about its 'guardrails' (restrictions), very high chance anthropic will censor it.

> chinese models are less censored Tried spicy geopolitical dispute questions?

Yes, and they aren't a problem. I get tired of people claiming this. Go download Qwen 3.6 and run it yourself and fire away.
Post reply on HN