All frontier models have been trained without any regards for IP protection laws. I don't see how anyone can argue in good faith that distillation is not fair game and does not ultimately "benefit humanity™"
Training a model is very expensive and creates something no individual rights-holder could. Distilling a model copies this value add and captures it without bearing the cost that created it.
Writing fingerprint analysis of responses reveals Kimi's similarity to Claude
31–40 of 73 posts
Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude
#32All frontier models have been trained without any regards for IP protection laws. I don't see how anyone can argue in good faith that distillation is not fair game and does not ultimately "benefit humanity™"
Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude
#33K3-to-Fable is blue at 0.42. Is 0.42 meaningful, or did we set 0.4 as the lower bound because it makes 0.42 look significant?
Sol-to-Fable is 0.69. It's dark yellow, making this look VERY different from 0.42. But is it? What do these numbers mean in absolute terms?
Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude
#34Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude
#35All frontier models have been trained without any regards for IP protection laws. I don't see how anyone can argue in good faith that distillation is not fair game and does not ultimately "benefit humanity™"
Does distilling actually violate any IP protection laws? Sure, it’s against their terms of service.
Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude
#36So this shows distance relative to other models, but I don't have a good sense for what these numbers say in absolute terms. K3-to-Fable is blue at 0.42. Is 0.42 meaningful, or did we set 0.4 as the lower bound because it makes 0.42 look significant? Sol-to-Fable is 0.69. It's dark yellow, making this look VERY different from 0.42. But is it? What do these numbers mean in absolute terms?
If you use the Claude Code harness on two models you will probably get very similarly styled output. I would not be surprised if K3 "stole" a lot of the harness that was leaked.
Notice how the only rows that even come close to being gray is GPT-5.4+Mini, diverging even from other GPT models. Is this because it has a wholly different training set, or (more likely IMO), did it just have a system prompt that leads to a different style?
Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude
#37Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude
#38So this shows distance relative to other models, but I don't have a good sense for what these numbers say in absolute terms. K3-to-Fable is blue at 0.42. Is 0.42 meaningful, or did we set 0.4 as the lower bound because it makes 0.42 look significant? Sol-to-Fable is 0.69. It's dark yellow, making this look VERY different from 0.42. But is it? What do these numbers mean in absolute terms?
Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude
#39All frontier models have been trained without any regards for IP protection laws. I don't see how anyone can argue in good faith that distillation is not fair game and does not ultimately "benefit humanity™"
It's a bit interesting how the open models are able to keep pace with the closed models except whole maintaining a steady following time.
Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude
#40So who's training on who's outputs?