Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

201–210 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#201

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

That is an apples to oranges comparison. An article about a video/book would have the relevant information in text form without needing to show the video "here is the new stuff shown in Apples 2 hour long WWDC keynote". If not is common that a comment in the discussion gives a summary as a tl;dr

With text articles behind paywalls the relevant information is hidden and only hinted at as a teaser.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#202
We developers like to pretend that LLM's are akin to humans and that they've been using things like NYTimes like humans as educational material.

But they are not. It's much simpler, proprietary writing is now integrated into the source code of OpenAI, it would be as if I would copy parts of other propriety code and copy paste it into my own codebase. Claiming copy paste is a natural evolving process of millions of years of evolution.

The fact that LLM's are so complicated and we don't know where it is, doesn't make it less so.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#203

Earlier quoted context omitted.

That's not the example. Here I proactively scrape NYT, summarise articles for a fee and sell that as a service. It's not people coming to me with some articles to summarise, and maybe then publishing it online. At some level it becomes a subversion of NYTs fees. First, say I subscribe and simply host the articles verbatim, for a fee. Clearly, that's not right. Suppose I change some spelling or word order, or use a sy…

>That's not the example. Here I proactively scrape NYT, summarise articles for a fee and sell that as a service. It's not people coming to me with some articles to summarise, and maybe then publishing it online. That's not what OpenAI is doing; it's not selling summarised articles as a service. Your example is a false equivalence. >This is kind of what LLMs do. And also feels like not fair use An LLM doesn't do this…

> An LLM doesn't do this unless you ask it to. And if you then take that output and publish it as your own, you're breaching the copyright, not OpenAI.

In this case, OpenAI is violating copyright by modifying, reproducing and distributing copyrighted content to its customer.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#204

We developers like to pretend that LLM's are akin to humans and that they've been using things like NYTimes like humans as educational material. But they are not. It's much simpler, proprietary writing is now integrated into the source code of OpenAI, it would be as if I would copy parts of other propriety code and copy paste it into my own codebase. Claiming copy paste is a natural evolving process of millions of ye…

> it would be as if I would copy parts of other propriety code and copy paste it into my own codebase.

It's not copy-pasted; it's compressed in a lossy manner. Even GPT4 has nowhere near enough memory to store the entirety of its training data in a non-lossy compression format. Just likes how humans compress the information we read.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#205

Earlier quoted context omitted.

Wow they want to kill it. I wonder if we've just lived through the golden Napster era of LLMs.

Just train on NYT articles no longer in copyright. We may be better for it.

So that would mean articles from the 1920s, provided that the authors of those articles have been dead for 70 years, or longer in some other countries.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#206

Earlier quoted context omitted.

I disagree. The verbatim part is the problem. You’re drawing a comparison to how humans operate except we’re not allowed to operate like that. While harder to do as a human, if memorised a copyrighted book and then did a live reading on TV, or produced replicas from memory and sold them (the most comparable example), I’d be sued. Humans produce derivative work all the time, and it’s fine for LLM’s to do that, but you…

>or produced replicas from memory and sold them (the most comparable example), I’d be sued. This is not the most comparable example, because it's not what ChatGPT is doing. The most comparable example is if you were hired as a contractor and the employer asked you to write verbatim some copyright content you'd memorised. If the employer then published it, they'd be the one liable, not you. >Humans produce derivative…

> The most comparable example is if you were hired as a contractor and the employer asked you to write verbatim some copyright content you'd memorised. If the employer then published it, they'd be the one liable, not you.

No, you'd both be liable. You are not allowed to create copies of a copyrighted work, even from memory, for any commercial purpose. Making it public or not is irrelevant.

This is more obvious with spftware: if I copy a version of AutoCAD that my previous employer bought and sell it to another company, or even just use it for my current employer without showing it to anyone else, I am violating the copyright on that software, and I am liable. Even though obviously no "publishing" happened.

Similarly, if you hire a decorator to paint Mickey Mouse on the inside walls of your private kindergarten, the decorator is violating Disney's copyright just as much as you are, even if neither of you has made that public.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#208
post #201

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

That is an apples to oranges comparison. An article about a video/book would have the relevant information in text form without needing to show the video "here is the new stuff shown in Apples 2 hour long WWDC keynote". If not is common that a comment in the discussion gives a summary as a tl;dr With text articles behind paywalls the relevant information is hidden and only hinted at as a teaser.

To make it an apples to apples comparison, look at submissions where the link submitted is the retail link to the IP. For example, look at all the book link submissions on AMZN...

https://news.ycombinator.com/from?site=amazon.com

None of these have the Pirate Bay or Library Genesis or Anna's Archive or the equivalent as the top comment.

Compare that to...

https://news.ycombinator.com/from?site=nytimes.com

And almost all of these have an archived version as the top comment.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#209
post #198

Under existing condition an AI news site seems like a good investment idea. Its AI could read all relevant news sources and retell them and republish them in its own articles. It could even have its own AI editors and contributors. Cannot see how human news companies could compete.

>Cannot see how human news companies could compete. News ultimately comes from physical sources on the ground, which currently AI has no way of doing.

I am sure it could easily rephrase the articles to tell them without quoting any real or verifiable sources. Many human news companies often do it too.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#210
post #60

Earlier quoted context omitted.

> AI bros What (or whom) do you consider to be an "AI bro?" This sort of ad hominem generalization usually accompanies a weak argument.

Young males that wear Tensorflow branded muscle tank tops and drive Mitsubishi Eclipse convertibles with the vanity plate OVERFIT. They are everywhere these days.

https://i.imgur.com/4tF7q8M.jpg
Post reply on HN