Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

761–770 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#761

Earlier quoted context omitted.

It depends. Google built a product out of scraping content (Google Search). But what I'm saying is that answering the question does not allow you to deduce anything about your rights; that's what I mean by "not a good question".

It can allow you to deduce something, depending on the answer. If we want to establish whether scenario A is fair use or not, and we all agree that A is "worse" (regarding fair use status) than some other scenario B, then if we also agree that B is not fair use, A by definition isn't either. The opposite is not true, of course: B being fair use does not imply that A has to be as well. I find that kind of upper/lower…

The problem is multi-dimensional, so bounding logic like this isn’t necessarily useful.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#762

Earlier quoted context omitted.

What you described is entirely fair use, actually. Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost all of them are directly sourced from NYT and others. They'll say "so and so happened, according to The Times" (and usually link the article there)

> What you described is entirely fair use, actually Just like during the pandemic how everyone became an epidemiologist, suddenly everyone's a copyright lawyer. I'll just dispute your assertion by saying: 1. Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most defini…

>if a work is purely derivative of a source work

This is the weakest part of the case(s) against OpenAI. "Derivative work" is a legal term of art meaning a direct adaptation, like writing a screenplay of a book or translating a book into another language.

NYT has a stronger case than Sarah Silverman here because they can show actual 'memorized' text rather than just summarization, but given that those memorizations are a) an unintended failure mode of the training process, and b) from an older version of the model that has been updated to no longer regurgitate memorized text, it's not really clear how in current form GPT could possibly be considered a derivative work.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#763

Earlier quoted context omitted.

> What you described is entirely fair use, actually Just like during the pandemic how everyone became an epidemiologist, suddenly everyone's a copyright lawyer. I'll just dispute your assertion by saying: 1. Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most defini…

> if a work is purely derivative of a source work CliffNotes, Wikipedia, etc. have huge quantities of summarized copyrighted work.

Summarization generally isn't considered a derivative work.

https://en.wikipedia.org/wiki/Derivative_work

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#764
post #422

At some point the burden of carrying 100 year old copywriter/patent law will become so onerous a burden on the pace of progress that its enforcement will be antihuman.

It already is, but I don't think this is a good example. NYT has a legitimate case here. They own the material they publish, and GPT-4 is shown to be able to recall entire articles verbatim. That's a violation, clear as day. The thing about lawsuits is that you make dozens of claims, and the court can rule in favor of some of them, and against others. The question of "is LLM training fair use?" hasn't made it to a hi…

A photocopier can reproduce entire articles verbatim, yet no one calls for the destruction of all photocopiers. In fact, many legitimate legal uses of photocopiers to reproduce whole newspaper articles take place commonly by archivists, journalists, students, etc.

It is the specific use of article photocopies to circumvent the normal sale of newspapers that becomes illegal.. and even that is questionable. If I read the newspaper left out in a waiting room and it keeps me from buying that days paper, this is not a criminal act.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#765

Earlier quoted context omitted.

> Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. >Is that fair use? IANAL, but doesn't sound like it. If you pay someone to do the summarisation for you, then you publish the content and charge a fee for it, you're the one liable, not the person you paid to summarise it for you. Similarly if you ask GPT to do it…

That's not the example. Here I proactively scrape NYT, summarise articles for a fee and sell that as a service. It's not people coming to me with some articles to summarise, and maybe then publishing it online. At some level it becomes a subversion of NYTs fees. First, say I subscribe and simply host the articles verbatim, for a fee. Clearly, that's not right. Suppose I change some spelling or word order, or use a sy…

How far is this from what reddit does?

I read a NYT article, then summarize it into a link title for reddit. Reddit then republishes the summary to all of its users.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#767
post #625

Earlier quoted context omitted.

Any publisher can opt out of google. Publisher also have substantial control over titles and snippets shown in google, whether an article appears in google news, etc Paraphrasing is also known as cloning and is often a copyright violation

Copyright law doesn't mention opt outs or search engine snippet controls. It's not clear to me that robots.txt is the singular thing that makes Google legal. In US copyright law facts cannot be copyrighted, so copyright on factual content like newspaper articles is limited. Simply replacing a few words wouldn't work, but I am certain that GPT-4 is capable of paraphrasing factual content at a level that would not be c…

If I make a website that scrapes NYT and passes it back and forth through a machine translator, say, English -> Spanish -> English, then the content will be slightly modified. Is this legal to make money off of?

Seems like the legal answer is unclear but, like Napster, such a system seems like it would lose in court.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#768
post #703

Earlier quoted context omitted.

It is a good question with a simple answer: no.

It depends. Google built a product out of scraping content (Google Search). But what I'm saying is that answering the question does not allow you to deduce anything about your rights; that's what I mean by "not a good question".

The general answer is no. Fair use is a special carve out legally that has to be determined individually. If your product is something that regurgitates NYT articles while stripping NYT of their source of revenue, that’s got fair odds to not qualify as fair use.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#769
post #761

Earlier quoted context omitted.

It can allow you to deduce something, depending on the answer. If we want to establish whether scenario A is fair use or not, and we all agree that A is "worse" (regarding fair use status) than some other scenario B, then if we also agree that B is not fair use, A by definition isn't either. The opposite is not true, of course: B being fair use does not imply that A has to be as well. I find that kind of upper/lower…

The problem is multi-dimensional, so bounding logic like this isn’t necessarily useful.

I think it can still provide value if the actual scenario at hand is so complex and fraught that conversations about it end up mostly fruitless (as I think is the case here). At least it can provide you with some mental handholds and supports for where to start reasoning about the problem, which hopefully helps in finding some small agreements, or at the very least, mutual understanding of each other's positions.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#770

Earlier quoted context omitted.

Copyright infringement is not avoided by changing some text so it isn’t an exact clone of the source. Determining whether a work violates a copyright requires holistic consideration of the similarity of the work to the copyrighted material, the purpose of the work, and the work’s impact on the copyright holder. There is not an algorithm for this, cases are decided on by people. There are algorithms that could detect…

And you think that it would be impossible to train a model to avoid outputs that are substantially similar to training data?

I certainly don't think it's impossible, but I think it is hard problem that won't be solved in the immediate future, and creators of data used for training are right to seek to stop wide availability of LLMs that regurgitate information they worked hard to obtain.
Post reply on HN