Live data from Hacker News

US Copyright Office found AI companies breach copyright. Its boss was fired

theregister.com

371–380 of 410 posts

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#371

Earlier quoted context omitted.

It's a very consistently Silicon Valley mindset. Seems like almost every company that makes it big in tech, be it Facebook and Google monetizing our personal data, or Uber and Amazon trampling workers' rights, makes money by reducing people to objects that can be bought and sold, more than almost any other industry. No matter the company, all claimed prosocial intentions are just window dressing to convince us to be…

> That's also why I'm really not worried about the "AI singularity" folks. The hype is IMO blatantly unsubstantiated by the actual capabilities, but gets pushed anyway only because it speaks to this deep-seated faith held across the industry. "AI" is the culmination of an innate belief that people should be replaceable, fungible, perfectly obedient objects, and such a psychosis blinds decision-makers to its actual li…

It will when it inevitably hits their wallets. Be it via the public rejection of a lower quality product, or court orders. But both sentiments move slow, so we're in here for a while.

Even with NFTs it still was a full year+ of everyone trying to shill them out before the sentiment turned. Machine learning, meanwhile, is actually useful but is being shoved into every hole.

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#372
post #348

Earlier quoted context omitted.

> Previously, a Meta executive in charge of project management, Michael Clark, had testified that Meta allegedly modified torrenting settings "so that the smallest amount of seeding possible could occur," which seems to support authors' claims that some seeding occurred. And an internal message from Meta researcher Frank Zhang appeared to show that Meta allegedly tried to conceal the seeding by not using Facebook ser…

>Meta allegedly modified torrenting settings "so that the smallest amount of seeding possible could occur," >Meta allegedly tried to conceal the seeding by not using Facebook servers while downloading the dataset to "avoid" the "risk" of anyone "tracing back the seeder/downloader" from Facebook servers Sounds like they used a VPN, set the upload speed to 1kb/s and stopped after the download is done. If the average Jo…

> If the average Joe copied that setup there's 0% chance he'd get sued

Citation needed. RIAA used to just watch torrents and sent cease and desists to everyone who connected, whether for a minute or for months. It was very much a dragnet, and I highly doubt there was any nuance of "but Your Honor, I only seeded 1MB back so it's all good".

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#373
post #258

Earlier quoted context omitted.

If a human were to reproduce, from memory, a copyrighted work, that would be illegal as well, and multiple people have been sued over it, even doing it unintentionally. I'm not talking about learning. I'm talking about the complete reproduction of a copyrighted work. It doesn't matter how it happens.

>I'm not talking about learning. I'm talking about the complete reproduction of a copyrighted work. It doesn't matter how it happens. In that case I don't think there's anything controversial here? Nobody thinks that if you ask AI to reproduce something verbatim, that you should get a pass because it's AI. All the controversy in this thread seems to be around the training process and whether that breaks copyright law…

Whereas, my report showed they were breaking copyright before the training process. Meta was sued for what I said they'd be sued for, too.

Like Napster et al, their data sets make copies of hundreds of GB of copyrighted works without authors' permission. Ex: The Pile, Commons Crawl, Refined Web, Github Pages. Many copyrighted works on the Internet also have strict terms of use. Some have copyright licenses that say personal use only or non-commercial use.

So, like many prior cases, just posting what isn't yours on HughingFace is already infringement. Copying it from HF to your training cluster is also infringement. It's already illegal until we get laws like Singapore's that allow copyrighted works. Even they have a weakness in the access requirement which might require following terms of use or licenses in the sources.

Only safe routes are public domain, permissive code, and explicit licenses from copyright holders (or those with sub-license permissions).

So, what do you think about the argument that making copies of copyrighted works violates copyright law? That these data sets are themselves copyright violations?

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#374

Earlier quoted context omitted.

The vast trove of copyright work has to refer to training. ChatGPT is likely on the order of 5-10TB in size. (Yes, Terabyte). There are college kids with bigger "copyright collections" than that...

No. The paragraph as a whole refers to the "outputs" of vast troves of copyrighted work. Disk size is irrelevant. If you lossy-compress a copyrighted bitmap image to small JPEG image and then sell the JPEG image, it's still copyright infringement.

I won't say it's irrelevant. How much you use is part of fair use considerations. Their huge collections of copyrighted works make them look worse in legal analyses.

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#375
post #186

I wonder when general internet sentiment moved from pro-piracy to IP maximalism. Fascinating shift.

Not having massively overfunded corporations exploit artists is not IP minimalism. Private persons stealing something is seen as tiny evil. But big corporation exploiting everyone else is entirely different thing.

IP minimalism is IP minimalism, regardless of who owns the IP.

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#376

I have yet to see someone explain in detail how transformer model training works (showing they understand the technical nitty gritty and the overall architecture of transformers) and also layout a case for why it is clearly a violation of copyright. You can find lots of people talking about training, and you can find lots (way more) of people talking about AI training being a violation of copyright, but you can't fin…

I'm not sure I understand your question. It's reasonably clear that transformers get caught reproducing material that they have no right to. The kind of thing that would potentially result in a lawsuit if you did it by hand. It's less clear whether taking vast amounts of copyrighted material and using it to generate other things rises to the level of copyright violation or not. It's the kind of thing that people woul…

> It's the kind of thing that people would have prevented if it had occurred to them, by writing terms of use that explicitly forbid it.

The AI companies will likely be arguing that they don’t need a license, so any terms of use in the license are irrelevant.

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#377
post #122

Earlier quoted context omitted.

>I'm not sure I understand your question. It's reasonably clear that transformers get caught reproducing material that they have no right to. The kind of thing that would potentially result in a lawsuit if you did it by hand. Is that a problem with the tool, or the person using it? A photocopier can copy an entire book verbatim. Should that be illegal? Or is it the problem that the "training" process can produce a mo…

Let's start with I think a case that everyone agrees with. If I were to take an image, and compress it or encrypt it, and then show you data file, you would not be able to see the original copyrighted material anywhere in the data. But if you had the right computer program, you could use it to regenerate the original image flawlessly. I think most people would easily agree that distributing the encrypted file without…

The model is not compressed data, it’s the compression algorithm. The prompt is compressed data. When you feed it a prompt it produces the uncompressed result (usually with some loss). This is not an analogy by the way, it’s a mathematical equivalence.

You can try and argue that a compression algorithm is some kind of copy of the training data, but that’s an untested legal theory.

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#378

If AI companies in the US are penalized for this, then the effect on copyright holders will only be slowed until foriegn AI companies overtake them. In such cases the legal recourse will be much slower and significantly limited.

Access to copyrighted materials might make for slightly better-trained models the way that access to more powerful GPUs does. But I don't think it will accelerate foundational advances in the underlying technology. If anything, maybe having to compete under tight constraints means AI companies will have to innovate more, rather than merely push scale.

But AI is mostly scale and only a little bit innovation. It’s undergraduate maths and a whole lot of computing power and data. Not being able to train on data on the internet would be a significant handicap.

Re: US Copyright Office found AI companies breach copyright. Its boss was fired

#380
post #365

Earlier quoted context omitted.

In the world you’re proposing, you would also not be able to make word-for-word copies of Harry Potter books, because Harry Potter wouldn’t exist.

why not? people write fiction all the time and put it on the internet for free. in fact, i'd say there's significantly more unpaid fiction writing in the world than paid.

People don't copy amateur fiction they can find for free. They copy (or rather, make derivative works of) successful commercial content because it is successful and well known.
Post reply on HN