Live data from Hacker News

“We have information that Moonshot distilled Fable for the development of K3”

twitter.com

461–470 of 742 posts

Re: “We have information that Moonshot distilled Fable for the development of K3”

#461
post #137

Earlier quoted context omitted.

> on the other hand the complete dismissal of copyright by AI labs Courts keep ruling over and over that an LLM trained on copyrighted works qualifies as a transformative work and is therefore fair use. They don't have to dismiss copyright law, this has always been allowed. The only thing they get in trouble for is pirating the works to get their hands on them.

"Keep ruling over and over" is way too strong. There have maybe been two rulings, nothing nationally binding, and most of the litigation is still ongoing. In particular, last I checked OpenAI and Microsoft are still badly threatened by the NYT lawsuit: https://law.justia.com/cases/federal/district-courts/new-yor... https://www.cnet.com/tech/services-and-software/publishers-o... This will have to wait for the Supreme…

> "Keep ruling over and over" is way too strong. There have maybe been two rulings, nothing nationally binding, and most of the litigation is still ongoing.

yep

https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...

> This ambiguity has resulted in extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we understand how Fair Use applies to any one step in AI training, let alone all.

> In the three lower court decisions so far, one held Fair Use did not apply (Thomson v Ross), one held Fair Use could apply (Kadrey v Meta) with the court suggesting more evidence was needed on the fourth factor ‘harm to the market’, and the third case held Fair Use may apply to some AI. As Fair Use is dependent on the specific facts at issue, none of these cases help educate the market or the public as to the limits of Fair Use in AI contexts.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#463
Who will invest in generating data for the frontier of AI if their output will immediately be used to train a competing model? Forget China vs. US, this applies within-country too. After exhausting all the publicly-accessible data on the internet, the frontier labs started spending hundreds of millions of dollars to generate data across a variety of fields. Just look at Mercor doing $1.2bln/year with 90% coming from the top labs (https://www.theinformation.com/articles/mercors-fast-growth-...). That pushes ahead what AI can do in medicine, science, coding, and math. But if other companies are going to free ride on this investment, it doesn't make sense to continue. So AI will largely hit a wall, frozen at the current level and work will all shift to cheaper inference.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#464
post #387

Earlier quoted context omitted.

I don't think that's what it's saying at all. It's saying that there's a level of creativity in model creation that isn't present in distillation.

Yes, this is what I was getting at.

There is an even higher level of creativity in creating books, songs and all sorts of art used in model training though. That's your apparent blindspot.

There is no world in which me vacuuming the entirety of human knowledge to make a genai model is ok but hoovering my model answers is not. The hypocrisy is stunning and risible.

Now if you go and make a model based on purely synthetic data and not a single work made by humans, you would have a valid point.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#466

[flagged]

The responses will continue until distillation-posting improves.

If we keep hearing accusations about how someone distilled something from someone, it seems like one of the few reasonable responses.

If my friend keeps complaining about how the inside of his car is wet, I will probably keep telling him to close the windows when it's raining, even if he thinks that is a tiresome take, lacking in insight.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#467

Earlier quoted context omitted.

If turning other peoples copyrighted work into a model is transformative enough to be protected then so is distilling that model into a different, better, model.

The models were built using copyrighted works, so why can't models be built using other models?

Because model output is probably far closer to software or a licensed work which possibly has greater protections than it is to copyright. There is far less possibility of fair use, it might be protected by patents, license or reverse engineering laws.

In any case the laws are being written now, but I doubt these will have worse protection than software does, which has far better protections than copyright

Re: “We have information that Moonshot distilled Fable for the development of K3”

#468
post #67

Earlier quoted context omitted.

It looks like these frontier-model companies don't really monitor their systems. Like OpenAI not realizing that it is their own AI which is attacking HuggingFace.

How does that connect with @throwa356262's argument?

That they should be able to find distillation 'attacks' if they had enough observability.

Re: “We have information that Moonshot distilled Fable for the development of K3”

#469

Earlier quoted context omitted.

This is why I don't give a shit that this is happening. It's actually kind of funny to me.

Unless you’re from mainland China, you should.

Why?

Serious question.

I'm from the US, and I think it's hilarious.

Post reply on HN