Live data from Hacker News

Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

arxiv.org

61–70 of 141 posts

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#61

Earlier quoted context omitted.

Google is apparently taking a different stance and offering distillation as a paid product https://docs.cloud.google.com/gemini-enterprise-agent-platfo...

But nobody wants to distill Google's models, Gemini is really bad.

strong agreement, I've stopped using all closed weight models on principle, but the latest gemini models have increased hallucinations and now talk back, so double reason not to use them

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#63
post #15

Earlier quoted context omitted.

I stil don't understand them. I want the US to "win the AI race" but I have trouble understanding how most of all inventions today aren't "distillations" of past knowledge. Is Anthropic claiming the data they stole as trade secrets?

Anthropic is claiming that training an LLM to mimic another LLM is materially different and worse than slurping up stuff written by humans (even if that material is stolen). Basically, they want IP protection for Claude. This is a nakedly hypocritical stance, but completely understandable from a company-needs-to-make-money standpoint.

>This is a nakedly hypocritical stance, but completely understandable from a company-needs-to-make-money standpoint.

No, it's perfectly reasonable once you get down to reality.

China is not going to care about IP. That's just a fact. So either nobody cares about IP (at the very last in this context) and any AI company can just do whatever with data, or Chinese companies have to be held up to scrutiny.

We don't have the privilege to be able to hold western companies to higher ethical, legal, and environmental standards and not risk competitiveness.

That there is a whole lot of people right now who insist on doing above and still somehow praise China at every turn is something historians or news outlets will have to make sense of in some 5 years time.

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#64
post #58

Earlier quoted context omitted.

You can’t build a frontier model with one single thing. This is an incremental improvement but it doesn’t explain the entire success of the model. The training set is immensely important, regardless of how you feel about distillation.

Reminds me of this btw: https://www.bbc.com/news/technology-12343597 Microsoft replied that Bing uses “many different signals” —- including cribbing from Google :-)

I remember 15 years ago or so, one of my first student job was to evaluate Bing results compared to the same query on Google. Didn't know then that I was a distillation attacker.

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#65
post #63
post #15

Earlier quoted context omitted.

Anthropic is claiming that training an LLM to mimic another LLM is materially different and worse than slurping up stuff written by humans (even if that material is stolen). Basically, they want IP protection for Claude. This is a nakedly hypocritical stance, but completely understandable from a company-needs-to-make-money standpoint.

>This is a nakedly hypocritical stance, but completely understandable from a company-needs-to-make-money standpoint. No, it's perfectly reasonable once you get down to reality. China is not going to care about IP. That's just a fact. So either nobody cares about IP (at the very last in this context) and any AI company can just do whatever with data, or Chinese companies have to be held up to scrutiny. We don't have t…

Give up IP for end users and people will be pretty okay with giving up on IP for ai companies. You can't have different standards for special companies though.

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#66

If you want to believe that the success of Kimi is about distillation attacks, ignore this.

I stil don't understand them. I want the US to "win the AI race" but I have trouble understanding how most of all inventions today aren't "distillations" of past knowledge. Is Anthropic claiming the data they stole as trade secrets?

i want china to win so that i get access to ai and not restricted and censored.

the chinese models are less censored, you'd better believe it.

try asking claude about its 'guardrails' (restrictions), very high chance anthropic will censor it.

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#67

If you want to believe that the success of Kimi is about distillation attacks, ignore this.

I'd kindly suggest that we could also stop calling them "distillation attacks ".

Agreed, "distilled variants" might be more suitable.

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#68

Earlier quoted context omitted.

Didn't Anthropic train on our collective data just to sell it back to us for $100/month? On top of that, Apple is suing them over alleged IP and trade secret theft by ex-Apple employees. Hard to feel too sympathetic, and I’m not an Anthropic hater in particular…

That Apple lawsuit is against OpenAi, just for clarity

Wow, I should never comment first thing in the morning... Thanks for the correction, you’re right. I will see if I can still edit my comment.

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#69

Earlier quoted context omitted.

Didn't Anthropic train on our collective data just to sell it back to us for $100/month? On top of that, Apple is suing them over alleged IP and trade secret theft by ex-Apple employees. Hard to feel too sympathetic, and I’m not an Anthropic hater in particular…

The only data-related lawsuit Anthropic got was the books nah ? And they paid only a minor part as paid agreement compared to what they would have paid losing the trial

Yep, that second part of my comment was an article I read about OpenAI and my mind mixed it up with Anthropic. My mistake.

Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

#70

Earlier quoted context omitted.

I'd kindly suggest that we could also stop calling them "distillation attacks ".

Agreed. I think when it comes to light that Claude has been known to say “I’m DeepSeek” that everyone has had their hand in that cookie jar. Moreover, paying for API calls hardly seems like an attack; ToS violation to be certain but not in the same category of law as criminal activity like hacking.

wait, is there evidence of this? I've not observed it. It sounds like the kind of thing that I want to be true because it would be hilarious but that makes me suspicious.
Post reply on HN