Live data from Hacker News

Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

typebulb.com

41–50 of 73 posts

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#41
post #10

All frontier models have been trained without any regards for IP protection laws. I don't see how anyone can argue in good faith that distillation is not fair game and does not ultimately "benefit humanity™"

Training a model is very expensive and creates something no individual rights-holder could. Distilling a model copies this value add and captures it without bearing the cost that created it.

A lot of those arguments could be said about writing a book, or a decent forum guide.

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#44

I didn’t participate in the discussion yesterday because I find it implausible Fable was available long enough (3-4 weeks cumulative?) to get data and train and have it fundamentally affect. But I don’t grok training enough to know that’s silly. If my new prior is you can…that’s a pretty thin moat that’s essentially indefensible.

My guess is they used the other Anthropic models extensively for synthetic data generation. The top most similar models for K3 are (lower means more similar) Fable 5 -> 0.42 Opus 4.8 -> 0.45 Sonnet 5 -> 0.45 Opus 4.7 -> 0.46 Grok 4.3 -> 0.52 There's an obvious jump at Grok 4.3, and it would not surprise me that the similarity there is because Grok used Anthropic models for training too (you can get the similarity lis…

Note that until last generation all Chinese models were relatively small. This was mainly due to lack of training hardware, but as soon as the new Huawei NPUs started shipping some Chinese labs switched to larger models:

Deepseek: 670B to 1.6T

Moonshot: 1.1T to 2.8T

Xiaomi: 310B to 1T

The new models also use a different architecture so I would assume in tests they will look different from previous generations.

Z.ai (GLM) and MiniMax on the other hand have continued training existing models (with some surgical changes to improve long context memory) so they should score similarly in these tests.

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#45
post #31

Earlier quoted context omitted.

Training a model is very expensive and creates something no individual rights-holder could. Distilling a model copies this value add and captures it without bearing the cost that created it.

Yes. I bet Moonshot paid for API access as opposed to pirating like Anthropic did.

[deleted]

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#46
>character trigrams

Worthless. Any LLM output similarity metric that uses n-grams as a source might as well measure the average temperature on Mars surface, no matter how much lipstick you put on it. It just can't have enough certainty. There used to be an n-gram benchmark popular on Twitter that showed extreme similarity of grok-3-beta to gpt-4.5-preview, while these models were trained on new base ones, came out 2 weeks apart, and were unmistakably different. Results were wildly inconsistent run to run. It didn't stop the crowd believing its creator in that DeepSeek R1 was trained on o1-preview (which was obvious bullshit as well, they were as different as two models can be). It's amazing how you can put anything on the web and everybody will believe you without checking or even understanding of what they're looking at.

K3 was trained on Claude's outputs, though - it repeats Anthropic's prompt injections 1:1 in its reasoning, which you should know if you ever tinkered with both models long enough. Good for them.

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#47
post #37

I would gladly make a donation of Claude accounts to Kimi, to support more distillation.

On a more serious note:

If I am working on something not sensitive and afterwards upload my chat log to a public dataset, can open models use that data for training?

Keep in mind, I paid for access to model and half of the work is mine (the prompts). But can the model provider forbid me from publishing my result as open data?

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#49
post #14

Earlier quoted context omitted.

Training a model is very expensive and creates something no individual rights-holder could. Distilling a model copies this value add and captures it without bearing the cost that created it.

What of the costs for creating the data that was used to train the model being distilled?

Do you people seriously believe that Moonshot and the Chinese personal-cult-state dont possess and train on the same torrents?

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#50

>character trigrams Worthless. Any LLM output similarity metric that uses n-grams as a source might as well measure the average temperature on Mars surface, no matter how much lipstick you put on it. It just can't have enough certainty. There used to be an n-gram benchmark popular on Twitter that showed extreme similarity of grok-3-beta to gpt-4.5-preview, while these models were trained on new base ones, came out 2…

[deleted]
Post reply on HN