Live data from Hacker News

Japan’s government will not enforce copyrights on data used in AI training

technomancers.ai

381–390 of 426 posts

Re: Japan’s government will not enforce copyrights on data used in AI training

#381
post #238

Earlier quoted context omitted.

Sure, but the key part there is "some of". They're necessarily able to produce verbatim copies only of the most duplicated, most repeated, most cited works -- and it's precisely due to their popularity that they're the only things worth including verbatim. I'm not going to opine on what the legality of that should be, but it's essentially the material considered most "quotable" in different contexts. I'm quite sure t…

I’m not sure it matters much that the current model can’t reproduce Harry Potter verbatim. If it can do smaller more quoted works now, it’ll tackle larger more obscure things in the future. It’s just a matter of time until it can output large copyrighted works, meaning the question of what to do when that happens is pretty relevant right now.

No it won't, because reproducing works verbatim is basically the definition of overtraining a model. That's a bug, not a feature.

A lot of further progress is going to be made towards making models smaller and more efficient, and part of that is reducing overtraining (together with progress in other directions).

Reproducing Harry Potter is a bug, because it's learning stuff it doesn't need to. So to the contrary, "it's just a matter of time" until this stuff decreases.

Re: Japan’s government will not enforce copyrights on data used in AI training

#382
post #319
post #184

I testified to the US Copyright Office this morning on AI in their roundtable session on AI and music[1]. A good portion of the focus of this panel was on whether copyrighted inputs (in this case, sound recordings and musical compositions) being fed into AI models for training purposes could plausibly constitute a fair use under existing US copyright law. Some of the comments here are missing the context of the recen…

> transformativeness is not dispositive in and of itself in the context of a fair use defense Could you dumb this sentence down for me? I would guess it means that making a derived work, changing the original, makes no difference in whether reproducing the work (in altered form) is fair use. But that sounds well-established, I can't imagine that movies would suddenly be legal to distribute if you just distribute the…

Sure. In an infringement lawsuit involving a fair use defense, courts will apply the "four prong" test [1] to determine whether or not such use is indeed fair use under copyright law. The first of the four prongs, the "purpose and character" of the use, is also known as "transformativeness." The Goldsmith/Warhol ruling (to simplify) said that Warhol's changes to Goldsmith's photograph were not sufficiently "transformative" even though they contained new expression (adding orange color etc.) because the end result effectively competed with the original photograph and therefore did not qualify as a fair use.

Right, your backwards movie example would fail the fair use test too. Nothing's really added, there's no new expression, it competes with the original, etc.

[1]: https://fairuse.stanford.edu/overview/fair-use/four-factors/

Re: Japan’s government will not enforce copyrights on data used in AI training

#383
post #110

Earlier quoted context omitted.

> I'm just imagining these "ghost AI artists" who publish catalogues of music using the audible likeness of more prolific artists. Possibly so, but who is the market for such a catalogue? I don’t see how the artist is going to lose out. I’m reminded, though, if the episode of Mad Men where they want to get the Beatles as the soundtrack to an ad, but on finding out that the Beatles won’t do it they try to get some mus…

A New Zealand political party got into trouble for using a "sound alike" of "Lose Yourself" by Eminem in a political advertisement. They licensed the song from a music library but the court determined the song they licensed was too close to the original. https://en.wikipedia.org/wiki/Eight_Mile_Style_v_New_Zealand...

Doesn’t sound like a core market for Eminem. I mean, I find it hard to think of use cases where an artist will really lose out in business terms because someone chooses an AI version of their music over the original. In this case I suspect they wouldn’t have really licensed or tried to licence Lose Yourself.

Re: Japan’s government will not enforce copyrights on data used in AI training

#384
post #344
post #289

Earlier quoted context omitted.

Did you just argue against fair use in general? Everything you wrote applies to it as well. Copyright has limits, it's not like right holders get to determine what those are. It's a balance of rights of creators and users.

is the use really all that fair, when thousands, if not millions, of works and artists get their works repurposed into services with "subscriptions" and "usage tokens" and other kinds of monetization, while directly competing with and against those very artists? or will it take until some markets would get completely destroyed and overtaken while people get displaced, for people to wake up and go 'wait, was that real…

I think societies haven't really determined what is fair in these cases. To give you a counter example: Millions of young artists look at art museums and incorporate what they see and learn into their own art. Creators of the images they look at get nothing from the proceeds young artists end up earning later in their lives. Is this unfair?

Re: Japan’s government will not enforce copyrights on data used in AI training

#385

Earlier quoted context omitted.

Humans can memorize their training data too... aka see something and then produce a copy (code, drawing, music etc). The principles underlying how LLMs and humans learn isn't really that different... just different levels of loss/fuzziness.

And when humans do that they may also infringe on copyright.

[deleted]

Re: Japan’s government will not enforce copyrights on data used in AI training

#386

Earlier quoted context omitted.

Does your ability to write fantasy books absolutely depend on having read those fantasy books as a kid? Was gaining the ability to write your own fantasy books and profit from them your only motivation to read those fantasy books? After gaining the ability to write fantasy books thanks to having read them, can you now produce fantasy books at a qualitatively different speed, scale, and conditions than any of the auth…

They already paid when they bought the books. Why would they need to pay more?

Say I want to write a screenplay and produce the resulting film, for profit, but I am literally unable to have any ideas whatsoever unless I base them on book that I read. With this in mind, and with this sole motivation, I buy and read the whole collection of Brandon Sanderson's novels and create a screenplay based exclusively on their content, for I have no ideas nor experiences of my own. I already paid when I bought the books. Why would I need to pay Brandon Sanderson any more?

Re: Japan’s government will not enforce copyrights on data used in AI training

#387

Earlier quoted context omitted.

The liability would lay with the company using the LLM product. This could mean that many companies won’t want to take on the risk unless there is decent tooling around warnings of infringement and listing sources.

I think liability lies with the person who uses the product to violate copyright. The hosting / producing company didn’t violate copyright if I use their model to make Mickey Mouse pictures. I did.

That’s what I said, the user would be liable. The user could be a company or an individual.

Re: Japan’s government will not enforce copyrights on data used in AI training

#388
post #344
post #289

Earlier quoted context omitted.

Did you just argue against fair use in general? Everything you wrote applies to it as well. Copyright has limits, it's not like right holders get to determine what those are. It's a balance of rights of creators and users.

is the use really all that fair, when thousands, if not millions, of works and artists get their works repurposed into services with "subscriptions" and "usage tokens" and other kinds of monetization, while directly competing with and against those very artists? or will it take until some markets would get completely destroyed and overtaken while people get displaced, for people to wake up and go 'wait, was that real…

It is as equally fair or unfair as humans who do fan art based on other characters or styles that they have seen and studied.

What is the difference between a model creating an infringing work of Iron man when prompted to do so and https://old.reddit.com/r/marvelstudios/search?q=%2Bflair%3AF... ?

After all, aren't those images directly competing with the artists of Marvel Studios and the artists who properly license it to create derivative works ( https://www.designbyhumans.com/shop/marvel/ )?

If we are going to say that creating images out of models that are training on something, then shouldn't clearly infringing works in the fan art category be handled the same way as they are competing with the very artists and companies who are the rights holders for the images and likenesses of the content?

Re: Japan’s government will not enforce copyrights on data used in AI training

#389

Earlier quoted context omitted.

Does your ability to write fantasy books absolutely depend on having read those fantasy books as a kid? Was gaining the ability to write your own fantasy books and profit from them your only motivation to read those fantasy books? After gaining the ability to write fantasy books thanks to having read them, can you now produce fantasy books at a qualitatively different speed, scale, and conditions than any of the auth…

> If the answer to those three questions is "yes", then I would argue that yes, you absolutely should have to pay royalties to the authors. Copyright lobbies can't have their own cake and eat it too, if you want to enforce such a different way of thinking about copyright compared to the current one, that would destroy the current industry and rightly so.

I would appreciate if you addressed the questions as literally stated.

I could just as well argue that businesses developing or making use of LLMs can't have their cake and eat it too. If they want their computer programs to enjoy the same rights and prerogatives as human creators do, they should be ready to demonstrate that those models are truly moral agents, with their own lived experiences, and thus deserve the status of legal persons as of themselves subject to the same laws and obligations as human beings.

Re: Japan’s government will not enforce copyrights on data used in AI training

#390
post #261

Earlier quoted context omitted.

If I were the copyright holder of such work, I would argue that the LLM was trained on text, including my copyrighted work, and that if the system produced text that a reasonable person who reads poetry would identify as the copyrighted work, the burden is then logically on the LLM owner to prove the LLM didn't regurgitate a piece of text from something it previously ingested. I think a jury would side with my argume…

The issue isn't that a generator lets you evade copyright somehow; it doesn't. The output is not the issue. If I sit in paint and my assprint happens to perfectly duplicate a Picasso, that's unlikely to fly in court if I try to sell copies. Picasso painted it first. The point at issue here is that some people are arguing that the models themselves are like a giant collective copyright infringement, since they are in…

I see your point now.
Post reply on HN