Live data from Hacker News

US Copyright Office found AI companies breach copyright. Its boss was fired

theregister.com

121–130 of 410 posts

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#121

I have yet to see someone explain in detail how transformer model training works (showing they understand the technical nitty gritty and the overall architecture of transformers) and also layout a case for why it is clearly a violation of copyright. You can find lots of people talking about training, and you can find lots (way more) of people talking about AI training being a violation of copyright, but you can't fin…

I would also like to see such explanation, especially one that explains how it differ from regular transformers found in video codecs. Why is a lossy compression a clear violation of copyright, but not a generative AI?

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#122

I have yet to see someone explain in detail how transformer model training works (showing they understand the technical nitty gritty and the overall architecture of transformers) and also layout a case for why it is clearly a violation of copyright. You can find lots of people talking about training, and you can find lots (way more) of people talking about AI training being a violation of copyright, but you can't fin…

I'm not sure I understand your question. It's reasonably clear that transformers get caught reproducing material that they have no right to. The kind of thing that would potentially result in a lawsuit if you did it by hand. It's less clear whether taking vast amounts of copyrighted material and using it to generate other things rises to the level of copyright violation or not. It's the kind of thing that people woul…

>I'm not sure I understand your question. It's reasonably clear that transformers get caught reproducing material that they have no right to. The kind of thing that would potentially result in a lawsuit if you did it by hand.

Is that a problem with the tool, or the person using it? A photocopier can copy an entire book verbatim. Should that be illegal? Or is it the problem that the "training" process can produce a model that has the ability to reproduce copyrighted work? If so, what implication does that hold for human learning? Many people can recite an entire song's lyrics from scratch, and reproducing an entire song's lyrics verbatim is probably enough to be considered copyright infringement. Does that mean the process of a human listening to music counts as copyright infringement?

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#123

I wonder when general internet sentiment moved from pro-piracy to IP maximalism. Fascinating shift.

No hard data to back this up, but anecdotally I'd place the AI/copyright sentiment shift around mid-late 2022. DALL-E 2 experimentation (e.g: [0]) in early-mid 2022 seemed to just about sneak by unaffected, receiving similar positive/curious reception to previous trends (TalkToTransformer, ArtBreeder, GPT-3/AI Dungeon, etc.), but then Stable Diffusion bore the full brunt of "machine learning is theft" arguments.

[0]: https://x.com/xkcd/status/1552279517477183488

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#124
post #51

Well, firing someone for this is super weird. It seems like an attempt to censor an interpretation of the law that: 1. Criticizes a highly useful technology 2. Matches a potentially-outdated, strict interpretation of copyright law My opinion: I think using copyrighted data to train models for sure seems classically illegal. Despite that, Humans can read a book, get inspiration, and write a new book and not be litigat…

Thank you - a voice of sanity on this important topic. I understand people who create IP of any sort being upset that software might be able to recreate their IP or stuff adjacent to it without permission. It could be upsetting. But I don't understand how people jump to "Copyright Violation" for the fact of reading. Or even downloading in bulk. The Copyright controls, and has always controlled, creation and distribut…

> But I don't understand how people jump to "Copyright Violation" for the fact of reading.

The article specificaly talks about the creation and distribution of a work. Creation and distribution of a work alone is not a copyright violation. However, if you take in input from something you don't own, and genAI outputs something, it could be considered a copyright violation.

Let's make this clear; genAI is not a copyright issue by itself. However, gen AI becomes an issue when you are using as your source stuff you don't have the copyright or license to. So context here is important. If you see people jumping to copyright violation, it's not out of reading alone.

> "People should not be allowed to read the book I distributed online if I don't want them to."

This is already done. It's been done for decades. See any case where content is locked behind an account. Only select people can view the content. The license to use the site limits who or what can use things.

So it's odd you would use "insane" to describe this.

> "People should not be allowed to write Harry Potter fanfic in my writing style."

Yeah, fan fiction is generally not legal. However, there are some cases where fair use covers it. Most cases of fan fiction are allowed because the author allows it. But no, generally, fan fiction is illegal. This is well known in the fan fiction community. Obviously, if you don't distribute it, that's fine. But we aren't talking about non-distribution cases here.

> "People should not be allowed to get formal art training that involves going to museums and painting copies of famous paintings."

Same with fan fiction. If you replicate a copyrighted piece of art, if you distribute it, that's illegal. If you simply do it for practice, that's fine. But no, if you go around replicating a painting and distribute it, that's illegal.

Of course, technically speaking, none of this is what gen AI models are doing.

> We just will not get to a sensible societal place if the dialogue around these issues has such a low bar for understanding the mechanics

I agree. Personifying gen AI is useless. We should stick to the technical aspects of what it's doing, rather than trying to pretend it's doing human things when it's 100% not doing that in any capacity. I mean, that's fine for the the layman, but anyone with any ounce of technical skill knows that's not true.

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#125

I have yet to see someone explain in detail how transformer model training works (showing they understand the technical nitty gritty and the overall architecture of transformers) and also layout a case for why it is clearly a violation of copyright. You can find lots of people talking about training, and you can find lots (way more) of people talking about AI training being a violation of copyright, but you can't fin…

I'm not sure I understand your question. It's reasonably clear that transformers get caught reproducing material that they have no right to. The kind of thing that would potentially result in a lawsuit if you did it by hand. It's less clear whether taking vast amounts of copyrighted material and using it to generate other things rises to the level of copyright violation or not. It's the kind of thing that people woul…

My comment is about training models, not model inference.

Most artists can readily violate copyright, that doesn't me we block them from seeing copyright.

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#126

Big Tech: We shouldn’t pay, each individual piece of content is worth basically nothing. Also Big Tech: We added 300.000.000 users worth of GTM because we trained in the 10 specific anime movies of Studio Ghibli and are selling their style.

The funny thing is that style is not copyrightable.

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#127
Intellectual property law is quickly becoming an institution of hegemonic corporate litigation of the spreading of ideas.

If it's illegal to know the entire contents of a book it is arbitrary to what degree you are able to codify that knowing itself into symbols.

If judges are permitted to rule here it is not about reproduction of commercial goods but about control of humanity's collective understanding.

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#128
post #51

Well, firing someone for this is super weird. It seems like an attempt to censor an interpretation of the law that: 1. Criticizes a highly useful technology 2. Matches a potentially-outdated, strict interpretation of copyright law My opinion: I think using copyrighted data to train models for sure seems classically illegal. Despite that, Humans can read a book, get inspiration, and write a new book and not be litigat…

The law covers these cases pretty well, it is just that the law has very powerful extremely rich adversaries, whose greed has gotten the better of them again and again. They could use work released sufficiently long ago to be legally available, or they could take work released as creative commons, or they could run a lookup, to make sure to never output verbatim copies of input or outputs, that are within a certain string editing distance, depending on output length, or they could have paid people to reach out to all the people, whose work they are infringing upon. But they didn't do any of that, of course, because they think they are above the law.

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#129
post #6

The released draft report seems merely to be a litany of copyright holder complaints repeated verbatim, with little depth of reasoning to support the conclusions it makes.

The required reasoning is not very deep though: If an AI reads 100 scientific papers and churns out a new one, it is plagiarism. If a savant has perfect recall, remembers text perfectly and rearranges that text to create a marginally new text, he'd be sued for breach of copyright. Only large corporations get away with it.

Is reading and memorizing a copyrighted text a breach of copyright? I.e. is creating a copy of the text in your mind a breach of copyright or fair fair use? Is it a breach of copyright if a digital “mind” similarly memorizes copyrighted text? Or is it only a breach of copyright to output and publish that memorized text?

What about loosely memorizing the gist of a copyrighted text. Is that a breach or fair use? What if a machine does something similar?

This falls under a rather murky area of the law that is not well defined.

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#130

Earlier quoted context omitted.

We are talking about the rights of the humans training the models and the humans using the models to create new things. Copyright only comes into play on publication. It's only concerned about publication of the models and publication of works. The machine itself doesn't have agency to publish anything at this point.

It's not only publication, otherwise people wouldn't be able to be successfully sued for downloading and consuming copyrighted content, it would only be the uploaders who get into trouble.

Do you have any links to cases where people were sued for downloading and consuming content without also uploading (eg, bittorent), hosting, sharing the copyrighted works, etc?
Post reply on HN