Live data from Hacker News

Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

typebulb.com

21–30 of 73 posts

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#21

Stolen data was stolen. Oh no! Anyway.

Can you elaborate? This looks very similar to the claim that distilling a model from Anthropic is the same thing as Anthropic distilling the model from information on the internet. Which is very flawed, since distillation requires the thing to exist in the thing it’s distilled from. And no LLM model existed in the information Anthropic used to train the model. Instead the model was built using information and utilizi…

Think of it in terms of distilled knowledge, not distilled LLMs.

I find both claims unsound, though. Knowledge or model behavior itself is not copyrightable, so all these claims just boil down to the "I am not happy with that" argument. You cannot claim someone is stealing something you don't own in the first place.

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#22
post #11

This looks to be more behavior based, not logit based? What is the actual claim?

It seems to boil down to China bad, US good, but with tech people speculating so it's definitely valid

honestly 99% of hn, reddit, etc. is the exact opposite of that statement.

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#23
post #18
post #14

Earlier quoted context omitted.

What of the costs for creating the data that was used to train the model being distilled?

At least 1.5B by stealing books, per a recent ruling. I don’t feel sorry for the model companies

They only had to pay for storing the books on a server for later possible use. They did not have to pay anything for the training which was declared fair use.

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#24
This data sort of disqualifies itself: unless Moonshot has a time machine, K3 should be more similar to Opus 4.5-4.8 than Fable 5.

Keep in mind, Anthropic started limiting access and introduced anti-distillation measures around 4.5-4.6 (?). So the majority of distillation should have happened on earlier models.

Maybe a better explanation is that they have access to the same training datasets? Which if private can again raise questions about theft, but on a very different level.

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#25

Stolen data was stolen. Oh no! Anyway.

Can you elaborate? This looks very similar to the claim that distilling a model from Anthropic is the same thing as Anthropic distilling the model from information on the internet. Which is very flawed, since distillation requires the thing to exist in the thing it’s distilled from. And no LLM model existed in the information Anthropic used to train the model. Instead the model was built using information and utilizi…

You might as well say a photocopy of a book was "built using information and utilizing new technology including hardware, software, transformer architecture, etc."

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#27
post #10

All frontier models have been trained without any regards for IP protection laws. I don't see how anyone can argue in good faith that distillation is not fair game and does not ultimately "benefit humanity™"

Training a model is very expensive and creates something no individual rights-holder could. Distilling a model copies this value add and captures it without bearing the cost that created it.

the outputs of the model have no property protections, and training a model on the outputs of another model does the exact same thing - its expensive and creates new value over what was in the input - a set of documents.

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#28

I didn’t participate in the discussion yesterday because I find it implausible Fable was available long enough (3-4 weeks cumulative?) to get data and train and have it fundamentally affect. But I don’t grok training enough to know that’s silly. If my new prior is you can…that’s a pretty thin moat that’s essentially indefensible.

My guess is they used the other Anthropic models extensively for synthetic data generation. The top most similar models for K3 are (lower means more similar)

Fable 5 -> 0.42

Opus 4.8 -> 0.45

Sonnet 5 -> 0.45

Opus 4.7 -> 0.46

Grok 4.3 -> 0.52

There's an obvious jump at Grok 4.3, and it would not surprise me that the similarity there is because Grok used Anthropic models for training too (you can get the similarity list for Grok and it does look like the top most similar models for Grok are either Anthropic models or some Chinese models).

The damning evidence that K3 used Anthropic models for training is that K3 is more similar to those models than it is to K2.6. If you look at the Anthropic, OpenAI or Google models, they are most similar with their own other models. Not so with K3, where K2.6 is less similar than 15 other models.

Now, why is Fable 5 the most similar to K3 and not Opus 4.8. I think it's quite likely that K3 did some fine tuning at the end, when Fable 5 became available. They probably had all the infrastructure in place, and Mythos had been announced for months, so they were probably waiting for the second the newest Anthropic model was released to start using it for synthetic data generation.

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#30
post #10

All frontier models have been trained without any regards for IP protection laws. I don't see how anyone can argue in good faith that distillation is not fair game and does not ultimately "benefit humanity™"

Training a model is very expensive and creates something no individual rights-holder could. Distilling a model copies this value add and captures it without bearing the cost that created it.

Could you give an example of the value that only training a model can create but none of the rights-holder could? I feel like if you got a direct, instant communication channel to any of the rights-holder that created the content in the training set of those models, you'd get more value than what the LLM could ever give you on any specific subject.
Post reply on HN