Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

401–410 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#401

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

I would be "happier" to pay a subscription to an aggregation platforms like hackernews or reddit to access archived articles that are linked to these sites. In turn a proportion of that could be passed on to the underlying publishers that I actually visit. I have nearly zero interest in reading articles that aren't linked to from an aggregation site. I don't want to read theguardian.com, or nytimes.com, or washington…

I would be happier to pay a small fee per article I want to read. But the norm seems a monthly subscription.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#402

Earlier quoted context omitted.

Which content making businesses earn humorous profit margins? Are all the journalist layoffs a fever dream? This is one of the more profitable ones, and only because they employ unscrupulous tactics: https://www.macrotrends.net/stocks/charts/NWS/news/profit-ma... This is NYT, the most successful news business: https://www.macrotrends.net/stocks/charts/NYT/new-york-times... As for movies/tv show/music makers, let’s ju…

The movie/tv show and music business can keel over and die tomorrow - it wouldn’t affect the value of art produced by humans at all. I see those more as exploitative leeches than as contributing anything positive. If only piracy would actually harm these businesses but alas as often demonstrated it has zero effect on their bottom line, if anything it increases their profits.

What do you mean by "art"?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#403
post #265

Earlier quoted context omitted.

To me, there is a sense that the news, which is real information about the society that we currently live in, should be availabe to all participants of that society. The notion of being a good citizen requires that one stays informed. Books, movies, videogames etc. don't have that role and are more consumption goods.

> should be available to all participants of that society. Who pays?

Everyone, if you don’t…

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#404

The lawsuit itself (which arstechnica links to): https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20... From page 30 and onwards has some fairly clear examples on how ChatGPT has an (internal) copy of copyrighted material which it will recite verbatim. Essentially if you copy a lot of copyrighted material into a blob and then apply some sort of destructive compression to it. How destructive would that compre…

> Essentially if you copy a lot of copyrighted material into a blob and then apply some sort of destructive compression to it. How destructive would that compression have to be for the copyright no longer to hold? My guess it would have to be a lot.

I imagine the goal is closer to "enough that no one notices we stole it", either in a way that it's not easily discoverable or even when directly analyzed there's enough plausible deniability to scrape by.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#405

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

> Typically I can't take a personal "tier" of a product and charge 3rd parties for derivatives of it.

I think you’re confusing terms of service and copyright. IANAL but what you describe sounds exactly like fair use to me, irrespective of how much you are paying NYT.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#406

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

> why we feel it's OK to pirate news articles, but not other IP.

Once the NYT pays reparations for the Iraq war, I'll be the first to stop pirating it.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#407
post #378

Earlier quoted context omitted.

>making training them at scale legally perilous Loading data to which you have no rights over into your software is legally perilous, yes. It's as easy as simply asking for and receiving permission from the data's rightsholders (which might require exchange of coin) to make it not legally perilous.

Sounds expensive.

If you want to do things with other people's stuff, yes it can get expensive.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#408

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

This tendency at Hacker News are also much more of a threat to The New York Times than what Open AI is doing. Even the places like blogs/Reddit/social media submissions that summarize the article and post the relevant quotes. Unlike the summary of a movie, summarizing all of the relevant parts of a news article is extracting almost all the value from it, and giving it away for free.

And the vast majority of people read news for it's breaking content, not for its archived content from years before (and I say this as someone who has often recommended the latter, but has gotten very few people to do so). So giving people that free breaking content (either in its entirety like on Hacker News, or summaries like you see all over social media) is actually a direct competition to the news business in a way that training an LLM on an article from months/years back isn't.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#409
post #373

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

> Is that fair use? As always, the answer is.. "it depends". I guess it depends mostly on the jurisdiction that applies to you. "Fair use" can have rather different legal meaning (or not exist at all) in different countries.

Also “fair use” does not use/define precedent - each case is assessed individually which really can be a flip of the coin.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#410

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

Largely because "news" aka facts is not and should not be copyrightable, so while the style, and exact format of the article may be copyrightable, the facts contained within are not. This makes a news story copyright murky in the eyes of wider society unlike a clearly 100% creative work like a TV Show or Movie. Further the news themselves self cannibalize, how many stories are just rewrites of stories from other outl…

Creative works like books, TV shows and movies contain facts too.
Post reply on HN