Earlier quoted context omitted.
False dichotomy right? Are Chinese labs impressively innovating? Clearly. However this doesn’t rule out possible gains due to distillation. I don’t know the degree of the latter but both things could certainly be true.
Didn't Anthropic train on our collective data just to sell it back to us for $100/month? On top of that, Apple is suing them over alleged IP and trade secret theft by ex-Apple employees. Hard to feel too sympathetic, and I’m not an Anthropic hater in particular…
Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
51–60 of 141 posts
Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
#52If you want to believe that the success of Kimi is about distillation attacks, ignore this.
False dichotomy right? Are Chinese labs impressively innovating? Clearly. However this doesn’t rule out possible gains due to distillation. I don’t know the degree of the latter but both things could certainly be true.
Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
#53Earlier quoted context omitted.
This is actually a well-known phenomenon in ML, called "The Bitter Lesson". > One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are search and learning. The full essay is worth a read, it's pre…
[flagged]
Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
#54Does any expert in the field know whether it is really the case that this intelligence we are seeing with frontier models is an "emerging" phenomena, only coming up when the architecture is scaled? Like isn't it weird that the 1 million parameter model with the same architecture can't solve basic puzzles but suddenly the 1 trillion parameter can conjure up counter-examples for the Jacobian conjecture? It's unintuitiv…
Mote-Carlo is pretty useful still. Not sure if your statement holds
Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
#55If you want to believe that the success of Kimi is about distillation attacks, ignore this.
False dichotomy right? Are Chinese labs impressively innovating? Clearly. However this doesn’t rule out possible gains due to distillation. I don’t know the degree of the latter but both things could certainly be true.
Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
#56Earlier quoted context omitted.
This is actually a well-known phenomenon in ML, called "The Bitter Lesson". > One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are search and learning. The full essay is worth a read, it's pre…
[flagged]
Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
#57Earlier quoted context omitted.
I stil don't understand them. I want the US to "win the AI race" but I have trouble understanding how most of all inventions today aren't "distillations" of past knowledge. Is Anthropic claiming the data they stole as trade secrets?
Anthropic is claiming that training an LLM to mimic another LLM is materially different and worse than slurping up stuff written by humans (even if that material is stolen). Basically, they want IP protection for Claude. This is a nakedly hypocritical stance, but completely understandable from a company-needs-to-make-money standpoint.
Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
#58If you want to believe that the success of Kimi is about distillation attacks, ignore this.
You can’t build a frontier model with one single thing. This is an incremental improvement but it doesn’t explain the entire success of the model. The training set is immensely important, regardless of how you feel about distillation.
https://www.bbc.com/news/technology-12343597
Microsoft replied that Bing uses “many different signals” —- including cribbing from Google :-)
Re: Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
#59Earlier quoted context omitted.
Anthropic is claiming that training an LLM to mimic another LLM is materially different and worse than slurping up stuff written by humans (even if that material is stolen). Basically, they want IP protection for Claude. This is a nakedly hypocritical stance, but completely understandable from a company-needs-to-make-money standpoint.
How does anyone justify this? How can you argue this in good faith?
Realer answer: A combination of the above, plus group/bubble effect of all your coworkers saying the same thing. You as a group, conflate a bunch of concerns together (China, no-guardrails-AI, etc), decide that your group will be the responsible stewards of AI, and then anybody "stealing your work" appears dangerous - both to your livelihood and to the human race.