Live data from Hacker News

Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

typebulb.com

31–40 of 73 posts

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#31
post #10

All frontier models have been trained without any regards for IP protection laws. I don't see how anyone can argue in good faith that distillation is not fair game and does not ultimately "benefit humanity™"

Training a model is very expensive and creates something no individual rights-holder could. Distilling a model copies this value add and captures it without bearing the cost that created it.

Yes. I bet Moonshot paid for API access as opposed to pirating like Anthropic did.

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#32
post #10

All frontier models have been trained without any regards for IP protection laws. I don't see how anyone can argue in good faith that distillation is not fair game and does not ultimately "benefit humanity™"

One mentions distillation to intimate that one is in fact still the leader. It has zero to do with fairness.

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#33
So this shows distance relative to other models, but I don't have a good sense for what these numbers say in absolute terms.

K3-to-Fable is blue at 0.42. Is 0.42 meaningful, or did we set 0.4 as the lower bound because it makes 0.42 look significant?

Sol-to-Fable is 0.69. It's dark yellow, making this look VERY different from 0.42. But is it? What do these numbers mean in absolute terms?

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#35
post #10

All frontier models have been trained without any regards for IP protection laws. I don't see how anyone can argue in good faith that distillation is not fair game and does not ultimately "benefit humanity™"

Does distilling actually violate any IP protection laws? Sure, it’s against their terms of service.

I get what you're saying and two wrongs do not make a right but the irony, and why people are even talking about this, is that the thing being distilled clearly, and knowingly, violated copyright & terms across the entire internet.

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#36
post #33

So this shows distance relative to other models, but I don't have a good sense for what these numbers say in absolute terms. K3-to-Fable is blue at 0.42. Is 0.42 meaningful, or did we set 0.4 as the lower bound because it makes 0.42 look significant? Sol-to-Fable is 0.69. It's dark yellow, making this look VERY different from 0.42. But is it? What do these numbers mean in absolute terms?

I also suspect the questions asked matter a lot, and the system prompts matter a lot, because "the map is built from nothing but the words they choose" - so this is more a measure of linguistic style than anything.

If you use the Claude Code harness on two models you will probably get very similarly styled output. I would not be surprised if K3 "stole" a lot of the harness that was leaked.

Notice how the only rows that even come close to being gray is GPT-5.4+Mini, diverging even from other GPT models. Is this because it has a wholly different training set, or (more likely IMO), did it just have a system prompt that leads to a different style?

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#38
post #33

So this shows distance relative to other models, but I don't have a good sense for what these numbers say in absolute terms. K3-to-Fable is blue at 0.42. Is 0.42 meaningful, or did we set 0.4 as the lower bound because it makes 0.42 look significant? Sol-to-Fable is 0.69. It's dark yellow, making this look VERY different from 0.42. But is it? What do these numbers mean in absolute terms?

I think it needs some more explanation - Fable is 0.42 from itself apparently so Kimi K3 is basically indistinguishable?

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#39
post #10

All frontier models have been trained without any regards for IP protection laws. I don't see how anyone can argue in good faith that distillation is not fair game and does not ultimately "benefit humanity™"

For me, I don't really care about the theft aspect but when people are claiming that these open models are better value or going to overtake anthropic/openai models, the implication that the open models are training of distilled data means all the "progress" they are making is just mimiced from the closed models.

It's a bit interesting how the open models are able to keep pace with the closed models except whole maintaining a steady following time.

Post reply on HN