Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

211–220 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#211
post #190

Earlier quoted context omitted.

Isn't it totally normal to write articles / blog posts that effectively summarize, and often quote from, news articles?

My impression is that it’s not necessarily legal, but going after bloggers and proving damages based is just a huge waste of their time. OpenAI came by with their fat stack of funding and changed that.

It is legal. Fair use. People have been doing it for ages. Almost every article you've ever read has some fair use of another article, book or news item, etc.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#212

Earlier quoted context omitted.

Yes, but this then hits against learning/understanding and compression being fundamentally the same thing . I can't think of a better way to argue in favor of "it's fine if human does it, therefore it's fine if LLM does it", than from the "lossy compression" angle.

It's not okay for a human to pirate, plagiarize, violate IP rights and laws, etc. But I disagree with the underlying assumption that you can anthropomorphize LLMs. Gradient descent and backpropagation don't take place in the brain. LLMs "learn" in the same way that Excel sheets "learn". Humans are living beings with needs and rights. A person being able to legally squat in a home doesn't mean that a drone occupying p…

> Gradient descent and backpropagation don't take place in the brain.

Not exactly, no, but the 'neurons that fire together wire together' way of learning has a pretty similar effect.

> LLMs "learn" in the same way that Excel sheets "learn".

I've never seen an excel sheet do anything like backpropagation.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#213
post #190

Earlier quoted context omitted.

Isn't it totally normal to write articles / blog posts that effectively summarize, and often quote from, news articles?

My impression is that it’s not necessarily legal, but going after bloggers and proving damages based is just a huge waste of their time. OpenAI came by with their fat stack of funding and changed that.

What parent poster meant is that it is normal that news organisations reference each other and report/cite/rephrase each other reports. For example all other news papers reported about the Watergate scandal reported by Bernstein&Woodward in the Washington Post.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#214

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

Can you read all of NYT and other things, and answer others' questions based on your knowledge? I'd imagine you can. I'm afraid you can't sidestep the question whether an LLM is more like a person who's read a lot or an archive/index.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#215

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

If NYT was a HN startup the link to the archived version would be banned and dang would be slamming the ban hammer.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#216

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

I find “4nn4’$ 4rch1v3 dot ORG” actually way better than pirate bay for pirating knowledge. It’s amazing the amount of books that copyright laws prevent us from finding https://www.theatlantic.com/technology/archive/2012/03/the-m...

Sure. It's just curious to me that news article have a pirated knowledge link as the de facto top comment, but link submissions to, for example, books for sale on Amazon don't have a link to Anna's Archive or equivalent.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#217

Earlier quoted context omitted.

My impression is that it’s not necessarily legal, but going after bloggers and proving damages based is just a huge waste of their time. OpenAI came by with their fat stack of funding and changed that.

It is legal. Fair use. People have been doing it for ages. Almost every article you've ever read has some fair use of another article, book or news item, etc.

When it becomes a service where you make money but the source doesn’t is it still fair use?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#219
> “The tragedy of the Luddites is not the fact that they failed to stop industrialization so much as the way in which they failed. Human rebellion proved inadequate against the pull of technological advancement.”

https://www.newyorker.com/books/page-turner/rethinking-the-l...

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#220

I read about this in the Times today (and am surprised that it wasn't on HN already). My guess is that the court will likely find in the Times favor, because the legal system won't be able to understand how training works and because people are "scared" of AI. To me, reading a book, putting it in some storage system, and then recalling it to form future thoughts is fair use. It's what we all do all the time, and I th…

In general, if you perform copyrighted works you are doing copyright infringement. There are certain exceptions (personal use, education, very small fragments with proper attribution, maybe a few others) but whether you are reading it aloud from a book or performing it from memory makes no difference.

So, if you setup a service like ChatGPT but powered by humans responding real time to queries, and these humans would occasionally reproduce large chunks of NYT articles, they and the service itself would be liable for copyright infringement. Even if they were all reproducing these from memory.

Now, this is somewhat different from the discussion of whether training the model on the copyrighted data, even if it had effective protections from returning copies of it, constitutes copyright infringement in itself. I believe this is a somewhat novel legal question and I can think of no direct corollaries.

I certainly don't think we can just handwave and say "at some level, when a human reads a copyrighted work, they are doing the same thing", because we really don't know if that is true. Artifical neural networks certainly have no direct similarity with the neural networks in the brain as far as we can tell. And, even if they did, there is no reason to give a machine the same rights that a human has - certainly not until that machine can prove sentience.

Post reply on HN