Live data from Hacker News

AI is in danger of being swallowed up by copyright law

heathermeeker.com

181–190 of 705 posts

Re: AI is in danger of being swallowed up by copyright law

#181
post #119
post #12

maybe one country or another will outlaw generative ai or ai art or media synthesis or whatever it ends up being called, but presumably they'll be left behind by rapid cultural and technical development in whatever countries don't the cat is out of the bag, the worms are out of the can, the feathers have blown away in the wind these developments seem very likely to be central to programming, all other kinds of engine…

Japan has explicitly amended their copyright code to enable machine learning on copyrighted data.[0] [0] https://storialaw.jp/en/service/bigdata/bigdata-12

So has the EU as part of the digital single market changes in 2019 (The so-called Text and Data mining exceptions)

Re: AI is in danger of being swallowed up by copyright law

#183
post #153
post #121

Earlier quoted context omitted.

AI companies can ask for permission if they want to train their models on other people's works Do you ask for permission when you train your mind on copyrighted books? Or observe paintings? Or listen to music? Do you ask for permission when you get new ideas from HN that aren't your own? Humans are constantly ingesting gobs of "copyrighted" insights that they eventually remix into their own creations without necessar…

> “Do you ask for permission when you train your mind on copyrighted books?” I’m not able to read billions of books in less than an hour. Even if we agree that machine learning is like human learning, scale commonly matters in law.

I recently worked on information extraction from 10K documents. GPT-3 needs about 7 days of operation in batch mode on one thread. It takes 40..70s to read one single document and report the extracted data. One MINUTE per page.

But I think you meant GPT-3 has seen many books during training, not during inference. You should know that training on millions of books is not the only way GPT-3 learns. It is just the foundation of its knowledge.

GPT-3 learns "in-context", that means it can learn a new word or a new task at first sight. It just needs a description or a few examples. This is the most powerful feature of GPT-3 - in-context learning. And when it comes to ICL, it is much like humans - only sees a few examples, not millions of books.

> “Do you ask for permission when you train your mind on copyrighted books?”

The nature of ICL is that it happens at prediction time. So GPT-3 would have to explicitly be instructed to learn a specific skill. Should it reject instructions if they are sourced from copyrighted books?

Re: AI is in danger of being swallowed up by copyright law

#184
post #116

Earlier quoted context omitted.

It’s strange to me that there’s a lot of overlap between people who think AI training should require explicit consent for every piece of training data, and people who think copyright and patents are insanely restrictive in the music/movies/literature/software world. It’s also worrying that requiring consent to train an AI model will inevitably lead to requiring consent to make handmade art that’s a little too similar…

It's probably those same people who publish their code under open source licenses instead of giving it to the public domain. I don't understand why people cling so hard onto every worthless little bit of code they write while also sort of half giving it away to almost everyone for almost any purpose.

1. People who use copyleft licenses do it because they know that the end result will be using a good license forever.

2. People using non-copyleft license just do it because public domain seems to have a complicated legal status across the world.

Re: AI is in danger of being swallowed up by copyright law

#185
post #7
post #3

The idea that it is copyright infringement if you train a neural network on copyright data means Waymo, Bing, Google are all illegal. If you include any copyrighted information in your web crawler neural network or if your training data for your autonomous software includes pictures of billboards or t-shirts or anything in the real world that is copyrighted you are a copyright infringer.

I think it comes down to use. Web crawlers like Google are fine because they index the web and then the search engine directs users to the original source. If instead Google recycled all the content they crawled and hosted everything on google.com while scrubbing all attributions from the pages then they’d fall afoul of copyright law (specifically the moral rights [1]). [1] https://en.wikipedia.org/wiki/Moral_rights

Google has stolen data from pages and showed it without attribution on their search page: https://docs.house.gov/meetings/JU/JU05/20190716/109793/HHRG...

Re: AI is in danger of being swallowed up by copyright law

#186
post #116

There's no part of AI that is being swallowed up by copyright. AI companies can ask for permission if they want to train their models on other people's works. It's not that hard, various image hosting sites have already added an opt-in/opt-out toggle to their services. Sites might even get away with using this stuff as compensation for free hosting. The fact of the matter is that the AI companies don't want to ask fo…

It’s strange to me that there’s a lot of overlap between people who think AI training should require explicit consent for every piece of training data, and people who think copyright and patents are insanely restrictive in the music/movies/literature/software world. It’s also worrying that requiring consent to train an AI model will inevitably lead to requiring consent to make handmade art that’s a little too similar…

What is strange in it exactly?

I can’t speak for everyone, but personally I find that copyright can be used properly or abused, at both sides (holder/consumer). It doesn’t mean that copyright is bad, only particular caregories of claims and usage are. But abusing copyrighted material from millions of little creators at insanely automated scale is another level of evil, especially when they explicitly require consent for exactly this type of use.

worrying that requiring consent to train an AI model will inevitably lead to requiring consent to make handmade art that’s a little too similar to some other existing artwork

That’s the root of misunderstanding, afaict. We can agree that at-scale processing is bad and that fair use is still okay. A human with a pen (or a text editor) can’t damage copyright at scale by learning terabytes of material in few weeks and producing the same amount in hours, so they can be excluded from this. Humans who use AI can, so they’re a target.

Re: AI is in danger of being swallowed up by copyright law

#187

There's no part of AI that is being swallowed up by copyright. AI companies can ask for permission if they want to train their models on other people's works. It's not that hard, various image hosting sites have already added an opt-in/opt-out toggle to their services. Sites might even get away with using this stuff as compensation for free hosting. The fact of the matter is that the AI companies don't want to ask fo…

Better title: The advancement of AI is being slowed by copyright. But "eating" is a fun word.

Even better title: tech giants don't want to pay copyright to small creators, but will easily be coherced to pay it to disney.

Re: AI is in danger of being swallowed up by copyright law

#188

There's no part of AI that is being swallowed up by copyright. AI companies can ask for permission if they want to train their models on other people's works. It's not that hard, various image hosting sites have already added an opt-in/opt-out toggle to their services. Sites might even get away with using this stuff as compensation for free hosting. The fact of the matter is that the AI companies don't want to ask fo…

[dead]

Re: AI is in danger of being swallowed up by copyright law

#189
post #119

Earlier quoted context omitted.

Japan has explicitly amended their copyright code to enable machine learning on copyrighted data.[0] [0] https://storialaw.jp/en/service/bigdata/bigdata-12

So has the EU as part of the digital single market changes in 2019 (The so-called Text and Data mining exceptions)

Not really.

Well they allow mining the data but nothing is said about the copyright of the collage output.

Re: AI is in danger of being swallowed up by copyright law

#190
post #119

Earlier quoted context omitted.

Japan has explicitly amended their copyright code to enable machine learning on copyrighted data.[0] [0] https://storialaw.jp/en/service/bigdata/bigdata-12

Very interesting! This would give Japan a real competitive productivity advantage if ML tools are banned or severely hampered elsewhere. It would also mean that other countries would need to ban not just scraping but also the resulting code or AI-generated media.

I'm sure chatgpt would be of much less interest worldwide if it only spoke japanese.
Post reply on HN