It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
I would be "happier" to pay a subscription to an aggregation platforms like hackernews or reddit to access archived articles that are linked to these sites. In turn a proportion of that could be passed on to the underlying publishers that I actually visit. I have nearly zero interest in reading articles that aren't linked to from an aggregation site. I don't want to read theguardian.com, or nytimes.com, or washington…
NY Times copyright suit wants OpenAI to delete all GPT instances
401–410 of 921 posts
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#402Earlier quoted context omitted.
Which content making businesses earn humorous profit margins? Are all the journalist layoffs a fever dream? This is one of the more profitable ones, and only because they employ unscrupulous tactics: https://www.macrotrends.net/stocks/charts/NWS/news/profit-ma... This is NYT, the most successful news business: https://www.macrotrends.net/stocks/charts/NYT/new-york-times... As for movies/tv show/music makers, let’s ju…
The movie/tv show and music business can keel over and die tomorrow - it wouldn’t affect the value of art produced by humans at all. I see those more as exploitative leeches than as contributing anything positive. If only piracy would actually harm these businesses but alas as often demonstrated it has zero effect on their bottom line, if anything it increases their profits.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#403Earlier quoted context omitted.
To me, there is a sense that the news, which is real information about the society that we currently live in, should be availabe to all participants of that society. The notion of being a good citizen requires that one stays informed. Books, movies, videogames etc. don't have that role and are more consumption goods.
> should be available to all participants of that society. Who pays?
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#404The lawsuit itself (which arstechnica links to): https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20... From page 30 and onwards has some fairly clear examples on how ChatGPT has an (internal) copy of copyrighted material which it will recite verbatim. Essentially if you copy a lot of copyrighted material into a blob and then apply some sort of destructive compression to it. How destructive would that compre…
I imagine the goal is closer to "enough that no one notices we stole it", either in a way that it's not easily discoverable or even when directly analyzed there's enough plausible deniability to scrape by.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#405If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…
I think you’re confusing terms of service and copyright. IANAL but what you describe sounds exactly like fair use to me, irrespective of how much you are paying NYT.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#406It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
Once the NYT pays reparations for the Iraq war, I'll be the first to stop pirating it.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#407Earlier quoted context omitted.
>making training them at scale legally perilous Loading data to which you have no rights over into your software is legally perilous, yes. It's as easy as simply asking for and receiving permission from the data's rightsholders (which might require exchange of coin) to make it not legally perilous.
Sounds expensive.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#408It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
And the vast majority of people read news for it's breaking content, not for its archived content from years before (and I say this as someone who has often recommended the latter, but has gotten very few people to do so). So giving people that free breaking content (either in its entirety like on Hacker News, or summaries like you see all over social media) is actually a direct competition to the news business in a way that training an LLM on an article from months/years back isn't.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#409If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…
> Is that fair use? As always, the answer is.. "it depends". I guess it depends mostly on the jurisdiction that applies to you. "Fair use" can have rather different legal meaning (or not exist at all) in different countries.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#410It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
Largely because "news" aka facts is not and should not be copyrightable, so while the style, and exact format of the article may be copyrightable, the facts contained within are not. This makes a news story copyright murky in the eyes of wider society unlike a clearly 100% creative work like a TV Show or Movie. Further the news themselves self cannibalize, how many stories are just rewrites of stories from other outl…