Live data from Hacker News

Extracting AI models from mobile apps

altayakkus.substack.com

121–130 of 250 posts

Re: Extracting AI models from mobile apps

#121
post #99

Earlier quoted context omitted.

The moment you earn money from it, that's not fair use anymore. When I last checked, unlimited access to said models were not free, plus it's not "research" anymore. - Addenda - For the interested parties, the law states the following [0]. Notwithstanding the provisions of sections 17 U.S.C. § 106 and 17 U.S.C. § 106A, the fair use of a copyrighted work, including such use by reproduction in copies or phonorecords or…

You can absolutely monetize works altered under fair use.

Any examples sans current AI models? I have not seen any, or failed to find any, to precise.

Re: Extracting AI models from mobile apps

#122

Earlier quoted context omitted.

Well obviously not in general. But when it comes to copyright law specifically, yes absolutely. That is the question I'm asking.

You're not going to get an answer you find agreeable, because you're hoping for an answer that allows you to continue to treat the tool as chattel, without conferring to it the excess baggage of being an individuated entity/laborer. You're either going to get: it's a technological, infinitely scalable process, and the training data should be considered what it is, which is intellectual property that should be being l…

Aren't animals a current example of a middle ground? They are incapable of authoring copyrightable works under current US law.

Re: Extracting AI models from mobile apps

#123
post #100

Earlier quoted context omitted.

The moment you earn money from it, that's not fair use anymore. When I last checked, unlimited access to said models were not free, plus it's not "research" anymore. - Addenda - For the interested parties, the law states the following [0]. Notwithstanding the provisions of sections 17 U.S.C. § 106 and 17 U.S.C. § 106A, the fair use of a copyrighted work, including such use by reproduction in copies or phonorecords or…

"Making money" does not immediately invalidate fair use, but it does wave a big red flag in the courts' faces.

So you say that, every law is a suggestion depending who's being tried?

Re: Extracting AI models from mobile apps

#124

Earlier quoted context omitted.

I don't see how this is a double standard. Comparing a person interacting with their culture is not comparable in any way. IMHO, it's kind of a wacky argument to make.

Can you elaborate on how it's not comparable? It seems obvious to me that it is -- they both learn and then create -- so what's the difference? If I can hire an employee who draws on knowledge they learned from copyrighted textbooks, why can't I hire an AI which draws on knowledge it learned from copyrighted textbooks? What makes that argument "wacky" in your eyes?

One is a collection of highly dithered data generated by machines paid for by a business in order to financially gain from the copyrighted works in order to replace any future need for copyrighted text books.

The other is a person learning from a copyrighted textbook in the legally protected manner, and whom and use the textbook was written for.

Re: Extracting AI models from mobile apps

#125

Earlier quoted context omitted.

You're not going to get an answer you find agreeable, because you're hoping for an answer that allows you to continue to treat the tool as chattel, without conferring to it the excess baggage of being an individuated entity/laborer. You're either going to get: it's a technological, infinitely scalable process, and the training data should be considered what it is, which is intellectual property that should be being l…

No, you’re missing the point of copyright. The point of copyright is to protect an exclusive right to copy, not the right to produce original works influenced by previous works. If an LLM produces original works that are influenced by the training data, that is not a violation of copyright. If it reproduces the training data verbatim, it is.

i weirdly agree with you, but also want to point out that “influenced by the training data” is doing some very heavy lifting there.

exactly how the new work is created is important when it comes to derivative works.

does it use a copy of the original work to create it, or a vague idea/memory of the original work’s composition?

when i make music it’s usually vague memories. i’d argue that LLMs have an encoded representation of the original work in their weights (along with all the other stuff).

but that’s the legal grey area bit. is the “mush” of model weights an encoded representation of works, or vague memories?

Re: Extracting AI models from mobile apps

#126

Earlier quoted context omitted.

And similarly, translating those sentences into data points is still a derivative work, like transcribing music and then making a new recording is still derivative.

derivative works still tend to be copyright violations.

Yes, that's what I'm saying. An LLM washing machine doesn't get rid of the copyright.

Re: Extracting AI models from mobile apps

#127
post #113

Earlier quoted context omitted.

If the resulting AI models are protected by copyright that invalidates the claim that AI models being trained on copyrighted materials is fair-use analogous to human beings becoming educated by exposure to copyrighted materials. Educated human beings are not protected by copyright, hence neither should trained AI models. Conversely, if a copyrightable work is produced based on work which itself is copyrighted, the re…

If I take 1000 books and count the distributions of the lengths of the words, and the covariance between the lengths of one word and the next word for each book, and how much this covariance matrix tends to vary across the different books, and other things like this, and publish these summaries, it seems fairly clear to me that this should count as fair use. (Such a model/statistical-summary, along with a dictionary,…

'Should the resulting work be protected by copyright? I’m not entirely sure…'

This has already been settled hasn't it? Don't companies have to introduce 'flaws' in order for data sets to be 'protected'? Just compiled lists of facts can't be protected. Which is why things like election result companies having to rely on NDAs and not copyright protections to protect their services on election night.

Re: Extracting AI models from mobile apps

#128
post #18

Earlier quoted context omitted.

> AI models are intellectual property If companies train on data they don't own and expect to own their model weights, that's hypocritical. Model weights shouldn't be copyrightable if the training data was pilfered. But this hasn't been tested because models are locked away in data centers as trade secrets. There's no opportunity to observe or copy them outside of using their outputs as synthetic data. On that subjec…

> If companies train on data they don't own and expect to own their model weights, that's hypocritical. Its not hypocritical to follow a line of legal analysis whoch holds that copying material in the course of training AI on it is outside the scope of copyright protection (as, e.g., fair use in the US), but that the model weights resulting from the training are protected by copyright. It maybe wrong, and it may be c…

The model weights are the result of an automated process, by definition, and thus not protected by copyright.

In my unusually well-informed on copyright but not a lawyer opinion, without any new legislation on the subject, I suspect that the most likely scenario for intellectual property rights surrounding AI is that using other people's works for training probably falls under fair use, since it's extremely transformative (an AI that makes text and a textual work are very different things) and it's extremely difficult to argue that the AI, as it exists today, directly impacts the value of the original work.

The list of what traing data to use is probably protected by copyright if hand-picked, otherwise just whatever web-crawler they wrote to gather it.

The AI models, as in, the inference and training applications are protected by copyright, like any other application.

The architecture of a particular AI model can be protected by patents.

The weights, as the result of an automated process, are probably not protected by copyright.

Re: Extracting AI models from mobile apps

#129

Earlier quoted context omitted.

You’re applying a double standard to LLM’s and human creators. Any human writer or artist or filmmaker or musician will be influenced by other people’s works, even while those works are still under copyright.

Human creators don't store that 'influence' in a digital machine accessible format generated directly from the copyrighted content though. Although with the 'good new everyone, we built the torment nexus' trajectory of AI my guess is at this point AI companies would just incorporate actual human brains instead of digital storage if that was the requirement.

Does that imply that if we invent brain upload technology, that my weights have every conflicting license and patent for everything I can quote or create? I don't like that precedent. I have complete rights over my noggin's contents. If I do quote a NYT article in it's entirely, that vould be infringement, but not copying my brain itself.

Re: Extracting AI models from mobile apps

#130
post #25

If I understand the position of major players in this field, downloading models in bulk and training a ML model on that corpus shouldn't violate anybody's IP.

IANAL But, this is not true it would be a piece of the software. If there is a copyright on the app itself it would extend to the model. Even models have licenses for example LLAMA is release under this license [1] [1] https://github.com/meta-llama/llama/blob/main/LICENSE

Yeah no.

An example for legal reference might be convolution reverb. Basically it's a way to record what a fancy reverb machines does (using copyrighted complex math algorithms) and cheaply recreate the reverb on my computer. It seems like companies can do this as long as they distribute protected reverbs separately from the commercial application. So Liquidsonics (https://www.liquidsonics.com/software/) sells reverb software but puts for free download the 'protected' convolution reverbs specifically the Bricasti ones in dispute (https://www.liquidsonics.com/fusion-ir/reverberate-3/)

Also, while a SQL server can be copyright protected, a SQL database is not given copyright protection/ownership to the SQL server software creators by extension of that.

Post reply on HN