Live data from Hacker News

Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

typebulb.com

61–70 of 73 posts

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#61
post #30

Earlier quoted context omitted.

Training a model is very expensive and creates something no individual rights-holder could. Distilling a model copies this value add and captures it without bearing the cost that created it.

Could you give an example of the value that only training a model can create but none of the rights-holder could? I feel like if you got a direct, instant communication channel to any of the rights-holder that created the content in the training set of those models, you'd get more value than what the LLM could ever give you on any specific subject.

>Could you give an example of the value that only training a model can create but none of the rights-holder could?

Yeah.

LLMs

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#62
post #31

Earlier quoted context omitted.

Training a model is very expensive and creates something no individual rights-holder could. Distilling a model copies this value add and captures it without bearing the cost that created it.

Yes. I bet Moonshot paid for API access as opposed to pirating like Anthropic did.

I understand you think Anthropic should have paid for the information it trained the models on. But im talking about all the costs to build a model. Do you think Anthropic didnt spend money to build these models? Did you not know that its actually very expensive?

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#63
post #21

Earlier quoted context omitted.

Can you elaborate? This looks very similar to the claim that distilling a model from Anthropic is the same thing as Anthropic distilling the model from information on the internet. Which is very flawed, since distillation requires the thing to exist in the thing it’s distilled from. And no LLM model existed in the information Anthropic used to train the model. Instead the model was built using information and utilizi…

Think of it in terms of distilled knowledge, not distilled LLMs. I find both claims unsound, though. Knowledge or model behavior itself is not copyrightable, so all these claims just boil down to the "I am not happy with that" argument. You cannot claim someone is stealing something you don't own in the first place.

You misunderstand the point. You can think both or either are morally right/wrong or good/bad for society. But distilling a model and building one using information online are fundamentally different. Even of you think the information the frontier models used was not fairly accessed.

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#64
post #25

Earlier quoted context omitted.

Can you elaborate? This looks very similar to the claim that distilling a model from Anthropic is the same thing as Anthropic distilling the model from information on the internet. Which is very flawed, since distillation requires the thing to exist in the thing it’s distilled from. And no LLM model existed in the information Anthropic used to train the model. Instead the model was built using information and utilizi…

You might as well say a photocopy of a book was "built using information and utilizing new technology including hardware, software, transformer architecture, etc."

Yeah but thats to my point. You can’t legally resell a photocopy of a book.

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#65
post #17

Earlier quoted context omitted.

Can you elaborate? This looks very similar to the claim that distilling a model from Anthropic is the same thing as Anthropic distilling the model from information on the internet. Which is very flawed, since distillation requires the thing to exist in the thing it’s distilled from. And no LLM model existed in the information Anthropic used to train the model. Instead the model was built using information and utilizi…

All ML/AI models are comparable to some form of compression(al beit lossy) of information and in this case copyrighted information. The OP is pointing to this as stolen data(by all the companies that started with pre trained models)

Models being compression algorithms doesn’t really make your point. Compression algorithms are copyrightable.

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#66
post #19

Earlier quoted context omitted.

Can you elaborate? This looks very similar to the claim that distilling a model from Anthropic is the same thing as Anthropic distilling the model from information on the internet. Which is very flawed, since distillation requires the thing to exist in the thing it’s distilled from. And no LLM model existed in the information Anthropic used to train the model. Instead the model was built using information and utilizi…

So what? If training on copyrighted data without authors consent is ok then distilling is ok as well.

Maybe both are ok. Maybe neither. Maybe one. The point is this does not follow

>If training on copyrighted data without authors consent is ok then distilling is ok as well.

Because building a model from distillation and building a model from raw data are not the same. You have to evaluate them independently. And legally its different as well. IP vs ToS (civil).

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#67
post #21

Earlier quoted context omitted.

Think of it in terms of distilled knowledge, not distilled LLMs. I find both claims unsound, though. Knowledge or model behavior itself is not copyrightable, so all these claims just boil down to the "I am not happy with that" argument. You cannot claim someone is stealing something you don't own in the first place.

You misunderstand the point. You can think both or either are morally right/wrong or good/bad for society. But distilling a model and building one using information online are fundamentally different. Even of you think the information the frontier models used was not fairly accessed.

I don't see any US labs suing their Chinese counterparts. It is practically impossible. That makes it a verbal battle, not a legal one.

From a strictly moral standpoint, it is illogical to state, "You stole things from my archive of stolen goods." LLM vendors need to morally own the knowledge before their accusations hold. It is unsound to claim ownership of something resulting from stolen property, regardless of the work you put into building it.

Either those accusations don't hold at all, all training is just "fair use", or the two processes you described are just different forms of "stealing."

Bottom line, even if we consider knowledge copyrightable, it is just stolen goods changing hands. You cannot claim ownership of a derivative work if you deny ownership to the sources your work was derived from. No one holds a higher moral standing than another.

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#68

Earlier quoted context omitted.

For me, I don't really care about the theft aspect but when people are claiming that these open models are better value or going to overtake anthropic/openai models, the implication that the open models are training of distilled data means all the "progress" they are making is just mimiced from the closed models. It's a bit interesting how the open models are able to keep pace with the closed models except whole main…

Distillation is absolutely not the reason they're good. It's not necessarily even done on a more capable model. It can even be done on itself and still bring improvement, or on a weaker model as well (see GLM and Gemini, which is definitely true because it repeats Deepmind's injections).

If distillation doesn't make them better, why do it?

If using Anthropic models for distillation doesn't make them better, why do it?

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#69
post #25

Earlier quoted context omitted.

You might as well say a photocopy of a book was "built using information and utilizing new technology including hardware, software, transformer architecture, etc."

Yeah but thats to my point. You can’t legally resell a photocopy of a book.

But you can use it to make an AI so then wouldn't distillation be fine?

Re: Writing fingerprint analysis of responses reveals Kimi's similarity to Claude

#70

Earlier quoted context omitted.

Distillation is absolutely not the reason they're good. It's not necessarily even done on a more capable model. It can even be done on itself and still bring improvement, or on a weaker model as well (see GLM and Gemini, which is definitely true because it repeats Deepmind's injections).

If distillation doesn't make them better, why do it? If using Anthropic models for distillation doesn't make them better, why do it?

It does, it's just not the reason. Most of the work is done before that point, and as I said z.ai used a weaker model. Besides, this is all strictly one-sided, as nobody knows how much Anthropic and OpenAI borrowed from Chinese labs' open research and weights (and they innovated a lot, to put it mildly, starting with first reasoning models worth talking about long before OAI did the same). Chinese labs are also severely restricted on hardware.

This entire story makes certain American AI shops look cartoonishly evil and Chinese ones relatively sane. Not only they want to grab without giving anything back, they also want to sabotage everyone else's AI research and do plenty of terrible things like media manipulation on the global scale and getting in bed with the government. This can't possibly end well, for the Americans in the first place.

Post reply on HN