Live data from Hacker News

Artificial Intelligence and Copyright: Request for comments

federalregister.gov

301–310 of 321 posts

Re: Artificial Intelligence and Copyright: Request for comments

#301
post #233

Earlier quoted context omitted.

I think it's learning styles in a way that's at least partially analogous, because it comes out with things that are reasonably original and not in the training data. I'm sure an LLM can write you an essay like that for any artist you want, but I'm not all that convinced those are meaningful even with humans. > As to your hypothetical That's the thing, it's not a hypothetical, it's a past story from here on HN. Someo…

>I think it's learning styles in a way that's at least partially analogous, because it comes out with things that are reasonably original and not in the training data. I don't think that is evidence that what it is doing is "learning". >I'm sure an LLM can write you an essay like that for any artist you want, but I'm not all that convinced those are meaningful even with humans. Well, it wouldn't be reflective of what…

> I don't think that is evidence that what it is doing is "learning".

When I say learning I mean something like "gaining new ability by studying how others did the same task, resulting in being able to produce novel output." I'm not quite sure what you are using the word to mean here, though I might agree that there are differences between what AIs do and what humans do, the question being what they are and whether they're important here.

I don't claim to know anything about the internal experience (if any) of an LLM writing such an essay and I can't really reason about that because I've never been an LLM, whereas I can at least relate to human experience. I think your assertion that it "wouldn't be reflective of what the LLM thinks" is a bit like saying that you don't think submarines are actually "swimming," as the saying goes, though. It may not "think" in human terms as we do, but it's certainly doing some kind of calculation that produces an equivalent output, so I have a lot of questions about whether we can say that on principle. We're well past passing the Turing test for a lot of things, either the original or censored form, these questions are getting less academic by the day.

> You say derivative but without any reference to what it actually means

We're talking about copyright law, so the meaning of derivative was borrowed from that, i.e. that AI model was producing works that could be reasonably thought to have infringed on the copyright of that painting when prompted for "a girl with a pearl earring" and this was held up to mean that AIs are just regurgitating training data and are therefore implicitly missing something essential to being an artist or what have you and all their work should be considered derivative works of the training data as far as copyright law is concerned.

Meanwhile, I'm saying that I think the AI should be judged about like a human artist would be to argue against the people who seem to want to say that the AI can't take input from copyrighted things without all of its output being tainted forever. We have no such requirement for humans and I don't see why it makes sense to add this new restriction on AIs specifically.

> Sorry I don't read every single thread about copyright on HN?

I'm not faulting you for not knowing, I'm faulting myself for assuming too much context and just trying to explain what I had in my head when writing that so you could understand how I came to think that. Hopefully this lets you see where I'm coming from.

Re: Artificial Intelligence and Copyright: Request for comments

#302
post #272

Earlier quoted context omitted.

I don’t think it makes sense for both model builders and the model’s users to separately obtain licenses for the same works used in the training set. A model trained on several copyrighted data sources cannot somehow be used in a way depending on a subset of those sources. So all parameters of usage and compensation should be settled by contract between the model builder and copyrighted data supplier, before the copy…

> Or to put it simply: using copyrighted material to create a model would NOT be considered fair use. The more I think about it, the more something along these lines seems like it might be the right way to think about it. When you play a DVD, for example, you copy the bits off the DVD, into the memory of your DVD player, and onto your screen; this is all explicitly considered "fair use" copying. But if you then copie…

> When you, as the human watch the DVD, bits of it get copied into your brain; but you don't then copy the bits of your brain to millions of other people -- they each have to make their own copy.

If you ripped The Little Mermaid, redrew every frame to combine it with The Fresh Prince of Bell-Air and moved things around in scenes to make it look like Ariel is Will Smith responding to sit-com dialogue, then it'd be fair use, regardless of how many people you show this new version to.

Fair use isn't about how or why you're doing with something. The definitions for fair use are very clearly laid out at https://www.law.cornell.edu/uscode/text/17/107

Re: Artificial Intelligence and Copyright: Request for comments

#303
post #301

Earlier quoted context omitted.

>I think it's learning styles in a way that's at least partially analogous, because it comes out with things that are reasonably original and not in the training data. I don't think that is evidence that what it is doing is "learning". >I'm sure an LLM can write you an essay like that for any artist you want, but I'm not all that convinced those are meaningful even with humans. Well, it wouldn't be reflective of what…

> I don't think that is evidence that what it is doing is "learning". When I say learning I mean something like "gaining new ability by studying how others did the same task, resulting in being able to produce novel output." I'm not quite sure what you are using the word to mean here, though I might agree that there are differences between what AIs do and what humans do, the question being what they are and whether t…

>When I say learning I mean something like "gaining new ability by studying how others did the same task, resulting in being able to produce novel output." I'm not quite sure what you are using the word to mean here, though I might agree that there are differences between what AIs do and what humans do, the question being what they are and whether they're important here.

I think the dictionary definition is more than sufficient: "the acquisition of knowledge or skills through experience, study, or by being taught." This is what I mean by running with your own made up definition.

>I don't claim to know anything about the internal experience (if any) of an LLM writing such an essay and I can't really reason about that because I've never been an LLM, whereas I can at least relate to human experience. I think your assertion that it "wouldn't be reflective of what the LLM thinks" is a bit like saying that you don't think submarines are actually "swimming," as the saying goes, though. It may not "think" in human terms as we do, but it's certainly doing some kind of calculation that produces an equivalent output, so I have a lot of questions about whether we can say that on principle. We're well past passing the Turing test for a lot of things, either the original or censored form, these questions are getting less academic by the day.

You are the one redefining words like "think" and "experience" not me. I'm not playing that game at all. After all, you are the one that is equivocating these processes between humans and AI by coming up with your own, much more broad concoctions.

>We're talking about copyright law, so the meaning of derivative was borrowed from that, i.e. that AI model was producing works that could be reasonably thought to have infringed on the copyright of that painting when prompted for "a girl with a pearl earring" and this was held up to mean that AIs are just regurgitating training data and are therefore implicitly missing something essential to being an artist or what have you and all their work should be considered derivative works of the training data as far as copyright law is concerned.

I'm familiar with copyright law, I'm not sure you are. A work can be derivative in a number of ways, some are legal, some aren't. It's not a new thing that some uses by a machine can be infringing, and others, non-infringing. Why now must it be that machines should be analyzed the same as humans all of the sudden?

>Meanwhile, I'm saying that I think the AI should be judged about like a human artist would be to argue against the people who seem to want to say that the AI can't take input from copyrighted things without all of its output being tainted forever. We have no such requirement for humans and I don't see why it makes sense to add this new restriction on AIs specifically.

Yes, I understand that. But I asked why it should be judged as a human, and you are saying because it "learns". But that's only based upon your re-defining the concept of learning in order to make it inhuman. The only reasonable arguments I've seen that AI outputs should be copyrightable are based on them being a tool that an artist can use. What you are saying is just dressed up anthropomorphization.

Re: Artificial Intelligence and Copyright: Request for comments

#304
post #301

Earlier quoted context omitted.

> I don't think that is evidence that what it is doing is "learning". When I say learning I mean something like "gaining new ability by studying how others did the same task, resulting in being able to produce novel output." I'm not quite sure what you are using the word to mean here, though I might agree that there are differences between what AIs do and what humans do, the question being what they are and whether t…

>When I say learning I mean something like "gaining new ability by studying how others did the same task, resulting in being able to produce novel output." I'm not quite sure what you are using the word to mean here, though I might agree that there are differences between what AIs do and what humans do, the question being what they are and whether they're important here. I think the dictionary definition is more than…

> I think the dictionary definition is more than sufficient: "the acquisition of knowledge or skills through experience, study, or by being taught." This is what I mean by running with your own made up definition.

I mean, if a human looked at a bunch of art, essays, etc. and then was able to produce similar works, we'd normally consider that "learning." What word would you use for being able to reproduce Picasso (or whomever) by looking at a bunch of examples?

Also I don't think I have defined "think" or "experience" at all. But I'd point out that I don't see anything like a principled boundary around them or that we can point to something that humans do that AIs don't or can't do. It seems to fall back on something that looks like qualia or subjective internal experience and philosophy hasn't resolved that with respect to other humans... except by analogy. "I think the other humans are like me and I have subjective internal experience, so they probably have it to, rather than being p-zombies."

If you have a better answer to that, feel free to tell me, it'd be interesting.

> It's not a new thing that some uses by a machine can be infringing, and others, non-infringing. Why now must it be that machines should be analyzed the same as humans all of the sudden?

Sure, I'll agree that it's not even necessary to consider the works transformative or whatever.

FWIW, I don't think that AIs should be getting their own copyrights or anything like that, I'm just saying that the training data shouldn't forever taint the output no matter what's produced.

Re: Artificial Intelligence and Copyright: Request for comments

#305

Earlier quoted context omitted.

I slightly disagree, in that I think the person using the tool should bear the burden of copyright. I.e. if the model outputs something under copywrite it merely can't be republished. In this same way, i can use Photoshop on proprietary data but I can't necessarily sell the results.

I'm so torn. On one hand, what you suggest seems to be a nearly ideal balance between advancing scientific progress and legal liability. By placing the legal burden to publish generated works on the person actually trying to publish, it allows for a more nuanced legal approach (i.e. the difference between "there are similarities to this work, but it's murky" or "you %100 stole that work"). On the other hand, is the c…

> is the company running the model themselves not already publishing all of that work and profiting from it?

no, because the model is transformative enough that it cannot be said to be a derivative works of the training set.

The model is in essence a form of distilled information, extracted from the training set. Information cannot be copyrighted - only expressions can.

Therefore, a model producer should have the right to use any pre-existing work, in the same way a person can, to study and internally memorize and extract information.

The reproduction of any of the training set data constitutes a copyright violation, but this is not done by the owner of the model, but by an end user of the model.

Re: Artificial Intelligence and Copyright: Request for comments

#306
post #64

[flagged]

Interesting example. I ran your prompt five times through GPT-3.5. While it didn’t use the name Hogwarts in those cases, three of the results had the character going to either the Academy of Arcane Arts (twice) or the Royal Academy of Arcane Arts, which seems clearly modelled on Hogwarts. I then tried the prompt five times with GPT-4. None of the resulting stories had the main character going to study magic at a scho…

Maeve is one of the main characters in HBO's Westworld. Like all of the "hosts" (bots), Maeve is said to be in a loop representing her part in a story that repeats until she is reprogrammed or decommissioned.

https://www.springfieldspringfield.co.uk/view_episode_script... :

> The hosts are supposed to stay within their loops, stick to their scripts with minor improvisations.

Re: Artificial Intelligence and Copyright: Request for comments

#307

Earlier quoted context omitted.

Interesting example. I ran your prompt five times through GPT-3.5. While it didn’t use the name Hogwarts in those cases, three of the results had the character going to either the Academy of Arcane Arts (twice) or the Royal Academy of Arcane Arts, which seems clearly modelled on Hogwarts. I then tried the prompt five times with GPT-4. None of the resulting stories had the main character going to study magic at a scho…

Maeve is one of the main characters in HBO's Westworld. Like all of the "hosts" (bots), Maeve is said to be in a loop representing her part in a story that repeats until she is reprogrammed or decommissioned. https://www.springfieldspringfield.co.uk/view_episode_script... : > The hosts are supposed to stay within their loops, stick to their scripts with minor improvisations.

That sounds like the source! Thanks.

Re: Artificial Intelligence and Copyright: Request for comments

#308
post #293

Earlier quoted context omitted.

Its not the tool that is covered by fair use. It is the creation of the tool that is covered by fair use. Is the tool itself supposed to be a copyright violation or is it a tool facilitating copyright violation by producing violating output? The later is something that can be tested because we have processes to compare works of art for it. If it is shown that LLMs produce mostly infringing art then we can and should…

> It is the creation of the tool that is covered by fair use. Copyright doesn’t restrict creation of something, it restricts (mainly) commercial distribution. Research, education and journalism etc are largely unaffected, and would still be. That said, I believe that selling access to the tool to the public already violates the copyright of the rights holders, even if it doesn’t produce similar works of art. The copy…

"That said, I believe that selling access to the tool to the public already violates the copyright of the rights holders, even if it doesn’t produce similar works of art. The copyrighted works increased the value of the product (otherwise why would they use it?)."

So it is similar to how ISPs argue that they should get a cut of streaming services because they enable another product.

I think it is also relevant that more than half of the globe will just completely ignore any regulation and any artist in a country with regulation will just have to compete with ever more empowered artists using all ai has to offer.

Re: Artificial Intelligence and Copyright: Request for comments

#309
post #254

Earlier quoted context omitted.

But that problem is already solved. Copyright holders are already protected from (I.e. can legally prohibit) distribution of obvious copies, or clearly derivative works. Regardless of how they were produced by hand, copy machine, Photoshop or with a model. The new problem is that artists styles are being “stolen” by incorporating their copyrighted work into models without their permission. And that problem can easily…

This is probably a somewhat unpopular opinion on HN, but it is where many of the artists I work with are generally trying to get to. Consent, compensation, and credit.

> Consent, compensation, and credit.

I just want to quote you. Nothing I need to say. That’s it.

Re: Artificial Intelligence and Copyright: Request for comments

#310

Earlier quoted context omitted.

> I used words very carefully. Then I am happy to use your word if that clarifies things. Just replace everything that I said about "human input" with "human authorship". And my point is that there are many things that a human can do using AI art that have large amounts of "human authorship" beyond just the boring case of prompting midjourney with a dumb prompt. > It depends on the level of human authorship. Oh hey!…

Its clear from this post that you aren't at all here in good faith, just to be glib and purposely misconstrue other people's words. good luck with life.

[deleted]
Post reply on HN