Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

171–180 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#171
I see few people here bring this up, so let me:

The US constitution says, The Congress shall have Power

> To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries;

So the Congress's power to make copyright and patent laws is predicated on promotion of science and useful arts (I believe this actually means technology). In a sense, the OpenAI being the forefront of our AI technology advancement is crucial to the equation. To hinder the progress by copyright is, in my mind, unconstitutional.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#173
post #66

Earlier quoted context omitted.

Well yeah, copying a work and using it for its original expressive purpose isn’t fair use, no? You have to use it for a transformative purpose. Suppose I’m selling subscriptions to the New Jersey Times, a site which simply downloads New York Times articles and passes them through an autoencoder with some random noise. It serves the exact same purpose as the New York Times website, except I make the money. Is that fai…

If they could find a single person who in natural use (e.g. not as they were trying to gather data for this lawsuit) has ever actually used ChatGPT as a direct substitution for a NYT subscription, I'd support this lawsuit. But nobody would do that, because ChatGPT is a really shitty way to read NYT articles (it's stale, it can't reliably reproduce them, etc.). All that is valuable about it is the way that it transfor…

That’s nonsense piracy. I never intend to own a truck, so when I need to haul a little something I go to Home Depot and steal a Ford off the lot for an hour? What if I stole all your commits, plucked the hard lines out of the ceremony, and then launched an equivalent feature the same week as you did, but for a competing software company? Would you or your employer deserve to get paid for my use of the slice of your work that was specifically useful for me? Yeah, and then some extra for theft.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#174

Earlier quoted context omitted.

Yes, but this then hits against learning/understanding and compression being fundamentally the same thing . I can't think of a better way to argue in favor of "it's fine if human does it, therefore it's fine if LLM does it", than from the "lossy compression" angle.

It's not okay for a human to pirate, plagiarize, violate IP rights and laws, etc. But I disagree with the underlying assumption that you can anthropomorphize LLMs. Gradient descent and backpropagation don't take place in the brain. LLMs "learn" in the same way that Excel sheets "learn". Humans are living beings with needs and rights. A person being able to legally squat in a home doesn't mean that a drone occupying p…

also if I write and article and quote some "text like this" [1] then that's not plagerism, but if my arguement is that the underlying assumption that you can anthropomorphize LLMs. Gradient descent and backpropagation don't take place in the brain. LLMs "learn" in the same way that Excel sheets "learn". Well, that's plagiarism and it's not allowed and people will get peeved and my career might get damaged.

I await the HN ban with fear..

[1] I'm not even doing referencing - so I am surely an LLM.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#175

Earlier quoted context omitted.

Another factor to consider is that neural nets can function as lossy compression, which becomes extremely evident when using models that are overfit. Sometimes they're so overfit that the compression isn't even lossy, and the data is encoded verbatim in the NN.

Yes, but this then hits against learning/understanding and compression being fundamentally the same thing . I can't think of a better way to argue in favor of "it's fine if human does it, therefore it's fine if LLM does it", than from the "lossy compression" angle.

It's fine for a human to remember it. It's not fine for a human redistribute it for money (legally speaking). That's copyright infringement.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#176

Why hasn't the Times also sued the Internet Archive? They've tried to block both the Internet Archive [1] and Open AI [2] from archiving their site, but why have they only sued OAI and not IA? The fact that they haven't sued IA which has comparatively little money would seem to indicate that this is not about fair use per se, but simply about profit-seeking and the NYT is selecting targets with deep pockets like OAI/…

What's wrong with that? If I was the NY Time's lawyers that what I would advise. What would it serve to bankrupt the IA, they can't pay anyway? These are corporations enforcing their rights against one another. There is nothing wrong with profit seeking from your copyright. That's literally their entire business model...they publish copyrighted content which they sell for a subscription. OpenAI and others could easil…

> What would it serve to bankrupt the IA, they can't pay anyway?

It would serve the termination of the infringement.

My point is that the Times doesn't particular seem to care about infringement per se, they care about getting their slice of the cut from that infringement.

It's like if a video game company or a movie company only attempted to sue illegal downloaders who had a certain net worth.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#177
> To me, reading a book, putting it in some storage system, and then recalling it to form future thoughts is fair use. It's what we all do all the time, and I think that's exactly what training is.

If the AI can recall the text verbatim then it's not at all the same. When we read we are not able to reproduce the book from our memory. Even if a human could memorise an entire book it's not at all practical to reproduce the book from that. The current AIs are not learning "ideas", they are learning orders of words.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#178

Earlier quoted context omitted.

> Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. >Is that fair use? IANAL, but doesn't sound like it. If you pay someone to do the summarisation for you, then you publish the content and charge a fee for it, you're the one liable, not the person you paid to summarise it for you. Similarly if you ask GPT to do it…

That's not the example. Here I proactively scrape NYT, summarise articles for a fee and sell that as a service. It's not people coming to me with some articles to summarise, and maybe then publishing it online. At some level it becomes a subversion of NYTs fees. First, say I subscribe and simply host the articles verbatim, for a fee. Clearly, that's not right. Suppose I change some spelling or word order, or use a sy…

>That's not the example. Here I proactively scrape NYT, summarise articles for a fee and sell that as a service. It's not people coming to me with some articles to summarise, and maybe then publishing it online.

That's not what OpenAI is doing; it's not selling summarised articles as a service. Your example is a false equivalence.

>This is kind of what LLMs do. And also feels like not fair use

An LLM doesn't do this unless you ask it to. And if you then take that output and publish it as your own, you're breaching the copyright, not OpenAI.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#179

I see few people here bring this up, so let me: The US constitution says, The Congress shall have Power > To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries; So the Congress's power to make copyright and patent laws is predicated on promotion of science and useful arts (I believe this actually mean…

Current AI is useless without people writing the articles in the first place.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#180
It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim.

And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or any other form of IP, it is in fact very much _not_ the norm for the top-rated comment to be a Pirate Bay link.

I think that's something worth reflecting on, about why we feel it's OK to pirate news articles, but not other IP.

And the reason I bring this up, is that it seems like Open AI has the same attitude: scraping news articles is OK, or at worst a gray area, but what if they were also scraping, for example, Netflix content to use as part of their training set?

Post reply on HN