Live data from Hacker News

Artificial Intelligence and Copyright: Request for comments

federalregister.gov

231–240 of 321 posts

Re: Artificial Intelligence and Copyright: Request for comments

#231

Much debate has been had about how existing copyright law applies to AI models. But once you get past that and start asking about how copyright should apply to AI models (as the copyright office is here) the answer in my mind becomes clear. Copyright, as defined in the U.S. Constitution, exists "to promote the Progress of Science and useful Arts"[1]. I can think of no better modern example of "the Progress of Science…

If you're going to bring up the origin of copyright you also need to consider the world copyright was made for. Back then, there was only one type of copying machine: a printing press. These were huge, expensive machines that could only practically be operated by corporations. Copyright was invented to protect authors from those corporations.

Your point 1 is what concerns me. AI models seem too much like the printing presses of old. They are available mainly to corporations and authors need to be protected from them. Otherwise there will be no incentive for anyone to publish anything novel as they know the corporation will slurp it up and make it "better" with their better model.

Re: Artificial Intelligence and Copyright: Request for comments

#232

Earlier quoted context omitted.

Trying, though it's hard to get noticed. https://news.ycombinator.com/item?id=37346620 And I want to participate in the community here, not merely mention the thing I've built.

> though it's hard to get noticed. I'd suggest to start with at least a brief paragraph on what thenose is, what it's goals are etc. I read your post, and found myself reading the technical workings of something I didn't know anything about.

Oh, thank you. Basically AI training datasets have been knocked offline recently by DMCAs, and the goal is to bring them back online in a place that can't be knocked offline. The most popular training dataset was The Pile, hosted by The Eye: https://pile.eleuther.ai/

Notice the links now 404. We tried to make a drop-in replacement for those links. All they have to do is change the-eye.eu to thenose.cc in the urls.

Unfortunately there's not a lot of ways to get their attention to let them know this exists now. I'll try emailing the contact address but I imagine they receive lots of spam, so I was hoping to try to get noticed by people like yourself first. Maybe a direct email is still the best way, but there's no guarantee they'll even be willing to change the urls due to legal risks. For all they know I could be logging the IP address of everyone who downloads it and forwarding it to authorities. But I'm not, and it's a frustrating problem to try to solve. I just want to help AI flourish.

This also serves as a template for someone else to do the same thing, so at least there can be multiple mirrors.

Thank you again. The fact that you even took the time to look it over meant a lot. If you have any other ideas, I'd be interested to hear.

Re: Artificial Intelligence and Copyright: Request for comments

#233
post #49

Earlier quoted context omitted.

My problem with this is that artists learn by studying other artists, cutting that off because it's AI rather than focusing on whether the resulting work is derivative, seems more of a problem to me. It seems to me that an AI can be used for either original work or derivatives, proving that you can get derivatives out of it has always struck me as no different than commissioning a copy of someone's work from a human…

Can an AI express to you how van gogh affected it as an artist? I'm not sure that AI is "learning" the way we say humans are "learning," when humans learn and study art. Obviously there is no debate that you can input van gogh into a model and produce something van gogh-like as a result. But I've not seen anything that indicates that the AI is learning anything about van gogh at all. Perhaps it comes down to whether…

I think it's learning styles in a way that's at least partially analogous, because it comes out with things that are reasonably original and not in the training data.

I'm sure an LLM can write you an essay like that for any artist you want, but I'm not all that convinced those are meaningful even with humans.

> As to your hypothetical

That's the thing, it's not a hypothetical, it's a past story from here on HN. Someone did that, asking for copies of a famous painting (Girl with a Pearl Earring) and got highly derivative items out of the model and we had a debate over whether that even means anything, because that's both a simple description of the painting and the name of a famous work, so it makes it so it can be ambiguous whether you asked for "Girl with a Pearl Earring" or a girl with a pearl earring in the prompting.

I agree that it looks like copyright infringement whether it's done by a human or AI, though. I guess a lot of people missed the prior discussion on HN.

Re: Artificial Intelligence and Copyright: Request for comments

#234
Every day of my life I wake up feeling more and more detached from the world.

That people are even debating this is so incredibly stupid to me.

The only people that will benefit from more onerous copyright are the major corporations that already have lawyers lined up and ready to fight their battles ad infinitum. See https://news.ycombinator.com/item?id=37347528

The everyday person will not benefit from any decisions made by courts on this matter. We will get the absolute shittiest implementation of copyright possible, see all other industries where copyright plays a role.

Copyright by itself is such a clever trick to prevent the world from advancing.

They convinced you to fight each other over scraps while they violate your copyright behind closed doors and use it to further their own agendas. They wield copyright like a weapon to effectively silence and dominate entire industries, leaving the average human unable to even comprehend how to fight back.

Lawyers and Copyright. Without them we would be so much better off.

Re: Artificial Intelligence and Copyright: Request for comments

#236

I believe we first need to answer the question of whether the copyright of the AI model’s source text or images affects the output. My opinion — and note I’m a software engineer, not a lawyer — is that an AI, being a statistical model and not generally intelligent, should not be allowed to disregard the copyright of its source material. This would, I think, require the AI’s creator to secure a license for all of its…

I hope your opinion isn't shared by lawmakers. Copyright is a relic of the past, and it needs to be put out of its misery. Trying to (mis)apply copyright here would just lobotomize the US. Existing companies would just technically operate out of a saner jurisdiction, and we'd be handing other countries a golden opportunity to leapfrog the US.

Re: Artificial Intelligence and Copyright: Request for comments

#237
post #56

I believe we first need to answer the question of whether the copyright of the AI model’s source text or images affects the output. My opinion — and note I’m a software engineer, not a lawyer — is that an AI, being a statistical model and not generally intelligent, should not be allowed to disregard the copyright of its source material. This would, I think, require the AI’s creator to secure a license for all of its…

My opinion as a SWE who is dating a lawyer (joke, not a serious qualification but it does provide some insight): Generative models traverse and interpolate high dimensional state spaces. These state spaces are created from input data. I would argue people do the exact same thing - the first main difference is we can use novel inputs (e.g. we can use images or words to develop our music/temporal state spaces and vice…

> ... a SWE who is dating a lawyer

> I would argue people do the exact same thing

Perhaps a ménage à trois with a neuroscientist would change your view on this.

Re: Artificial Intelligence and Copyright: Request for comments

#238
post #96

Earlier quoted context omitted.

Is intelligence really a factor here? Say I use the same training set as one of these LLMs, copyright protected text and all, and use it to derive a compression algorithm that uses very little space to store tokens and token sequences that are common in that huge collection of text. The resulting compression scheme includes some sort of statistical artifact derived from that copyrighted text. Is that allowed? And if…

Very good question indeed. A lot of these questions are somewhat ethical/moral in nature. E.g. is it okay to take someone else's creative work, process it through some algorithm, to create a service like ChatGPT? Or a compression algorithm? I don't know. It's awesome to see the Copyright office request input from both sides of the argument.

It worries me that so much focus is on two sides that may not have the end-users' best interest much in mind. The companies building the models may have an incentive to regulate models to keep smaller players or open source projects away. Artists mostly seem totally anti any solutions as even laws that allow models trained on purely public domain art would be bad for them. If laws around this are shaped primarily by the wishes of those two groups I am not sure things will end up well at all for those of us that want the tools to keep improving and remain reasonably free (including applications you can install locally and run on your own GPU).

Re: Artificial Intelligence and Copyright: Request for comments

#239
post #32

There are three copyright issues here; datasets, model weights, and model outputs. Dataset copyright is pretty well defined and things can often be used under fair use. Fair use decisions are done with a four prong test and really decided by the courts on a case-by-case basis. Model weights cannot currently be copyrighted. They are the output of a mechanical process over the dataset. However, software faced a similar…

Realistically an AI model is basically just a very complicated piece of software. The model weights are akin to the software code, the model outputs are akin to the outputs a user of the software creates, and the datasets are akin to the intellectual property put into the software by the developer to create the code. In the same way that a developer could not simply steal someone elses intellectual property in order…

It's not stealing, and the term "intellectual property" should be put to rest:

https://www.gnu.org/philosophy/not-ipr.en.html

Your opinions on what should be copyrightable are wrong, and fortunately the courts agree.

Re: Artificial Intelligence and Copyright: Request for comments

#240
post #164

Earlier quoted context omitted.

> I can do it, but I cannot claim ownership over the characters. of course not. But you can claim ownership if you don't call those characters their original names, and make sufficient changes to the design (how sufficient is determined by a court of law - thus expenses). > DC or Marvel if I try to do this at scale. The show 'invincible'[1] has a character that is a basic copy of superman. And yet, you will find that…

> make sufficient changes to the design I think that’s one of the issue. The transformation done by these tools are mechanical. Even if it may be extensive. The human input is too small. Omniman may have similarities with Superman, but he is not him in the larger context of the story. LLMs can not yet be that consistent for marketable output that deserves to be copyrightable. I’m perfectly fine for LLMs to aid with s…

> The human input is too small.

That's a huge assumption, especially for image generation models.

Post reply on HN