Live data from Hacker News

AI is in danger of being swallowed up by copyright law

heathermeeker.com

351–360 of 705 posts

Re: AI is in danger of being swallowed up by copyright law

#351
post #119
post #12

maybe one country or another will outlaw generative ai or ai art or media synthesis or whatever it ends up being called, but presumably they'll be left behind by rapid cultural and technical development in whatever countries don't the cat is out of the bag, the worms are out of the can, the feathers have blown away in the wind these developments seem very likely to be central to programming, all other kinds of engine…

Japan has explicitly amended their copyright code to enable machine learning on copyrighted data.[0] [0] https://storialaw.jp/en/service/bigdata/bigdata-12

> although the laws of foreign countries have provisions with the same effect as Article 47-7 of Japan’s Copyright Act, all of which limit use to development for non-commercial purposes and development by research organizations

It reads like there's a bunch of countries that have similar legislation, interesting though.

Re: AI is in danger of being swallowed up by copyright law

#352
post #289
post #244

Earlier quoted context omitted.

> Do you ask for permission when you train your mind on copyrighted books? Yes, you do need to buy books, which gives you permission to read them.

This is literally what the AI does as well. It didn't walk into a bookstore and steal all the books off the shelf, it read through material made available to it entirely legally. The thing that authors are trying to argue here is that they should get to control what type of entity should be allowed to view the work they purchased. It's the same as going "you bought my book, but now that I know you're a communist, I t…

> It didn't walk into a bookstore and steal all the books off the shelf, it read through material made available to it entirely legally.

Github ignored the licenses of countless repos and simply took everything posted publicly for training. They didn't care whether it was available to them entirely legally, they just pretended that copyright doesn't exist for them.

Re: AI is in danger of being swallowed up by copyright law

#353
post #140

Earlier quoted context omitted.

So is market harm "Some courts have held this factor to be the most important in the analysis." https://ilt.eff.org/Copyright__Fair_Use.html#Market_Harm

I’m unclear about this. Let’s say a movie comes out and I make a YouTube review using brief clips or screenshots from the movie. Since my review is transformative, I should be in the clear (I think?). But when it comes to market harm, does the tone of my review effect the enforceability of copyright? As in, if my review is negative it would harm the market for people going to watch the movie vs a positive review righ…

I am not a lawyer but I'd think:

You're not directly competing with the movie though, your work is a review, not a feature film.

If you were to make a parody movie from the material of the movie itself, directly taking scenes and altering them to your liking but still relying on the viewer recognizing the original in it, you'd have a harder time, I think.

Re: AI is in danger of being swallowed up by copyright law

#354
post #314

Earlier quoted context omitted.

It is superior, but it is absolutely still learning. ML can do either facsimile or imitation far better than a human mind can. You seem to be conflating both things and suggesting that ML only does facsimile, which is where the potential legal problems are.

It is superior, but it is absolutely still learning No, it is not. Memorization =! understanding. I can teach a parrot to spew the times table, good luck getting it to understand how to apply it. And a parrot is billions upon billions of times more capable than any current AI algos.

You don't need to "understand" art to create it though. Good art, sure, maybe, but not really. Plenty of brilliant musicians who know fuck all about music theory, but they can crank out tunes.

Diffusion would appear to me to work in much the same way. It doesn't understand what's good ("works", creates acceptable output) or why, but it knows it when it sees it, and has the tools to refine it.

Re: AI is in danger of being swallowed up by copyright law

#355
post #289

Earlier quoted context omitted.

This is literally what the AI does as well. It didn't walk into a bookstore and steal all the books off the shelf, it read through material made available to it entirely legally. The thing that authors are trying to argue here is that they should get to control what type of entity should be allowed to view the work they purchased. It's the same as going "you bought my book, but now that I know you're a communist, I t…

> they should get to control what type of entity should be allowed to view the work they purchased No, that's not it. It's more like if I memorized a bunch of pop-songs, then performed a composition of my own whose second verse was a straight lift of a song by Madonna. I would owe her performance royalties. And I would be obliged to reproduce her copyright notice, so that my audience would know that if they pull the…

There are lots of people arguing against the training itself. And people arguing against all outputs, even when there is no detectable copying. I don't know how you missed those takes. You're arguing the wrong point here. Many people do want to say "no ai can look".

Re: AI is in danger of being swallowed up by copyright law

#356

Earlier quoted context omitted.

If you have an image, then train a neural network on that image, then use the neural network to reconstruct that image in detail, then the NN by definition contains enough information used to reconstruct that image - hence, a copy. With NNs trained on thousands or millions of data entries, this concept becomes fuzzy in the same way as you described - a short summary likely wouldn't be considered a copy, just like a 6…

The thing is, the “good” models can’t reconstruct the image in detail. It’s considered a sign of “overfitting” if you reconstruct the input exactly. Even if you put the exact query that was associated with that image, you’ll get the weighted average (feature-wise) image associated with the query. This applies to all like machine learning models without loss of generality.

That doesn't mean I can't recover the image (or at least get really close to one) using a different query, does it? Edit: It's nonlinear after all.

Re: AI is in danger of being swallowed up by copyright law

#358

Earlier quoted context omitted.

It gets called plagiarism, and there are lots of lawsuits preventing this.

It's not plagiarism at all. The AI is trained on 5 billion images yet it stores only 4gb of data. Thus it is impossible that it stores the actual work. For any image that the AI generates, you can't point to any image in the training data that the image is derived from.

There are some overfittings where you can point to source images, but that number is a lot smaller than 5 billion.

Re: AI is in danger of being swallowed up by copyright law

#359
post #116

There's no part of AI that is being swallowed up by copyright. AI companies can ask for permission if they want to train their models on other people's works. It's not that hard, various image hosting sites have already added an opt-in/opt-out toggle to their services. Sites might even get away with using this stuff as compensation for free hosting. The fact of the matter is that the AI companies don't want to ask fo…

It’s strange to me that there’s a lot of overlap between people who think AI training should require explicit consent for every piece of training data, and people who think copyright and patents are insanely restrictive in the music/movies/literature/software world. It’s also worrying that requiring consent to train an AI model will inevitably lead to requiring consent to make handmade art that’s a little too similar…

Not sure what you intended to imply, but I don't think those two consents are related that much to worry about. Copyright licenses usually are written down like this: "[you are allowed to] use, reproduce, modify, adapt, perform, display, distribute" and so on.

When a new technology is introduced, for example the compact disc was invented, lawyers get to poke whether that "distribute" applies to the music CDs, or just to vinyls and music tapes (because at the time of granting that license, CDs weren't yet a thing! gotcha!).

The answer to this conundrum might vary in different countries, and we can have fun discussing that in the context of AI, but it does not affect how handmade art shouldn't be too similar.

Re: AI is in danger of being swallowed up by copyright law

#360

There's no part of AI that is being swallowed up by copyright. AI companies can ask for permission if they want to train their models on other people's works. It's not that hard, various image hosting sites have already added an opt-in/opt-out toggle to their services. Sites might even get away with using this stuff as compensation for free hosting. The fact of the matter is that the AI companies don't want to ask fo…

How is it much different from a search index? it’s just a new interface to get at some info, rather than Google and Firefox, it already pre-browsed the web for you, and is displaying back the content. If the end user gleans some actual copyrighted work from the search they still need permission to use it, but it’s also likely it’s just a derivative, or the end user is just reading an example and learning from it at c…

> How is it much different from a search index?

A search index usually links to the source. Without that a search index is worthless, you can't use content if you don't even know where it comes from and who holds the rights.

Google search links to sources like Wikipedia in its info boxes, because without that you can't know whether the info is reliable or sourced from my brother's coworker's imaginary flat-earther friend.

Post reply on HN