Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

271–280 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#271

Earlier quoted context omitted.

The NYT argument is going to be that they put up a site, own the copyright for their content and make that content available for either a human to read it for themselves, or software to index for something commonly understood as a search engine. Those terms do not entitle the training of LLMs for commercial use. Therefore, cease and desist. Oh and destroy anything that was created by violating the terms of our licens…

Terms of Use are a thing, and if the Times can prove that OpenAI infringed their web terms by scraping, they may have a case... but terms of use probably won't monetize well or give them enough leverage to prevent OpenAI from using their data anyway and may end-up distracting from the main copyright suit.

Violating TOS, at least to scrape and use later, is legal.[0] I'm not sure how the ruling interacts with LLMs, but I'm sure OpenAI's lawyers would bring it up.

[0]: https://www.forbes.com/sites/zacharysmith/2022/04/18/scrapin...

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#272
post #255

Earlier quoted context omitted.

It would be great if such "creative" works were simply impossible to monetize. I already don't pay for these and try to find unknown artists / writers who have a job and do stuff simply because they enjoy doing it.

That’s… awful.

What's awful here? This way it's just a social interaction with both sides satisfied, where's the problem?

I see a problem today, being spammed with shitty commercial "works" whereas I'd like to see something genuine and not just made for money.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#273

Earlier quoted context omitted.

> They absolutely are copying in the course of training the model Only in the sense that a Cisco router is copying in the course of sending me the article, which we've all agreed doesn't count as infringement. The bigger problem is that the plaintiff has to show it is more likely than not that that sequence of words came from their text and not some other source , which is going to be obscenely difficult. > And ChatG…

> Only in the sense that a Cisco router is copying in the course of sending me the article, which we've all agreed doesn't count as infringement. But it does count as infringement, if its not explicltly or implicitly licensed (because it is necessary to a use that is licensed or necessary to a use that does not itself require a license but is the normal use for for which a licensed copy, which you have, is sold and u…

> OpenAI wpuld be on the hook, civilly and potentially criminally for any infringement

I don't think you've thought about this hard enough

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#274

Earlier quoted context omitted.

Copyright law includes many exceptions explicitly for libraries.

Copywrite law doesn't ban people from reading the source altogether. That is what is being proposed, ban AI from being allowed to 'read' the content. The real argument is how does a human brain aggregate knowledge and then profit from it, and is it really that different from an AI model aggregating knowledge. They both read in data, perform calculations on the data, and spit out something.

This is a religious belief and not scientific knowledge or strong legal argument.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#275

Earlier quoted context omitted.

they had to make a copy of the original to get it into their system in the first place!

In order to render that page, it probably was copied dozens of times all over my RAM. Do I owe NYT money now?

yes. they have a hard paywall!

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#276

Earlier quoted context omitted.

Terms of Use are a thing, and if the Times can prove that OpenAI infringed their web terms by scraping, they may have a case... but terms of use probably won't monetize well or give them enough leverage to prevent OpenAI from using their data anyway and may end-up distracting from the main copyright suit.

Violating TOS, at least to scrape and use later, is legal.[0] I'm not sure how the ruling interacts with LLMs, but I'm sure OpenAI's lawyers would bring it up. [0]: https://www.forbes.com/sites/zacharysmith/2022/04/18/scrapin...

FYI LinkedIn actually won that case after appealing once more: https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#277
post #221

Earlier quoted context omitted.

>which the press is now framing as “trying to hide the use of copyrighted data” Yea, now I can't read the paper and talk about it to other people it seems. The Right to Read was a prophecy I guess?

An individual or group of individuals doing this and sharing their views/summary vs. a profit-oriented program funded by major technology companies scraping this information and spitting it back out algorithmically does seem different to me. Yes, perhaps both things are on the same "sliding scale", but I do not view them as fundamentally equivalent actions.

So I can use an open source LLM like Llama then?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#279
post #272

Earlier quoted context omitted.

That’s… awful.

What's awful here? This way it's just a social interaction with both sides satisfied, where's the problem? I see a problem today, being spammed with shitty commercial "works" whereas I'd like to see something genuine and not just made for money.

If not for the ability to monetize, there would be a small fraction of the total available work out there. From music, to movies, to video games. And for many, the quality we come to enjoy just wouldn’t be possible. Do you think we’d have a Skyrim, or GTA, or equivalent if there weren’t millions to be made to employ thousands of people to make it happen? What about the largest and most influential films and TV shows of the past decade? These things take money to produce. I could see the argument for music, maybe, but even then I don’t think it’d be sustainable. Just because an artist may enjoy working on their art after working 50 hours a week to pay the bills doesn’t mean they should HAVE to if their art is good enough and desired enough to sustain them, thus allowing them to create more and of a higher caliber (in theory).

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#280

I think anybody should have the right to protect the word combinations they own by not publishing them on the internet.

Copyright law doesn’t stop applying because your website is accessible over the internet. I think generative AI is cool but training an AI to do the thing you do using your words is pretty clearly not allowed under current copyright law, and not only that but it makes people who use AI look like bad people who are fine stealing other people’s work for their own enjoyment.
Post reply on HN