Live data from Hacker News

“We have information that Moonshot distilled Fable for the development of K3”

twitter.com

381–390 of 742 posts

Re: “We have information that Moonshot distilled Fable for the development of K3”

#381

Earlier quoted context omitted.

Perhaps even more importantly, the current frontier LLM models are self-admittedly the product of enormous quantities of copyright infringement and even less savory inputs, so calling them out for distilling the fruit of that tainted tree reads as highly hypocritical at best.

No - distillation is not data inputs. Raw materials vs. Value add. They are different things, like ore and metal. Distillation is a new thing we need to understand, it's probably closer to IP than not.

Are you suggesting data input is further from IP than distillation?

That would stun me, but it's a little hard to read.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#386
post #265

Earlier quoted context omitted.

While I agree on a moral level, I think there is a distinction to be made. Training a SOTA model takes a huge amount of resources and expertise so the people doing the training are adding a lot of value along the way. I think this is much less true for distillation (which is kind of the whole point). ed: to clarify, I totally agree that a huge chunk of the value in LLMs is coming from the source material. My point wa…

I don't know man. This reads like "yeah we stole your grain, but making bread is hard ."

It sure is, but it doesn't matter. Whatever position that generates more economic activity is declared legal using some nonsense retconned logic "because we said so".

Re: “We have information that Moonshot distilled Fable for the development of K3”

#387

Earlier quoted context omitted.

I like this comment because its argument only makes sense if you assume that the entire world's output of books and art did not require a huge amount of resources and expertise to make, nor did it add any value. It's the most CS-major take ever!

I don't think that's what it's saying at all. It's saying that there's a level of creativity in model creation that isn't present in distillation.

Yes, this is what I was getting at.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#388

Earlier quoted context omitted.

> Distillation is impossible to stop Lots of things are impossible or very difficult to stop completely but measures can be taken to reduce their prevalence.

Sure, but the problem is that it hurts legit people too.

> the problem is that it hurts legit people too.

What hurts other people too?

Re: “We have information that Moonshot distilled Fable for the development of K3”

#389
post #265

Earlier quoted context omitted.

While I agree on a moral level, I think there is a distinction to be made. Training a SOTA model takes a huge amount of resources and expertise so the people doing the training are adding a lot of value along the way. I think this is much less true for distillation (which is kind of the whole point). ed: to clarify, I totally agree that a huge chunk of the value in LLMs is coming from the source material. My point wa…

I like this comment because its argument only makes sense if you assume that the entire world's output of books and art did not require a huge amount of resources and expertise to make, nor did it add any value. It's the most CS-major take ever!

If turning other peoples copyrighted work into a model is transformative enough to be protected then so is distilling that model into a different, better, model.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#390

Earlier quoted context omitted.

I like this comment because its argument only makes sense if you assume that the entire world's output of books and art did not require a huge amount of resources and expertise to make, nor did it add any value. It's the most CS-major take ever!

This is a misrepresentation though. The LLM output, is not the same as the input - there is value add. Of course works used as raw inputs to LLMs required work and are reasonably subject to IP concerns - but they are different. It's possible that the LLM makers 'owe' the content creators that created the content they used to make their products - it's an interesting but separate question. We could very well end up wh…

> but they are different.

How, and why?

> We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.

That is the current state of legal rulings - LLM output is public domain, not copyrightable.

Post reply on HN