Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

511–520 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#511

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

As a supporter of piracy in the general case, I tend to agree with your observations, including that pirating NYT (FT, NPR, ...) articles is somehow some kind of different class of offense as, say, stealing a movie or mp3.

(Books, to me, are separate still, in that I like to have a physical copy (and generally see the authors as humans who deserve compensation, rather than mega-orgs that deserve eternal torment), so I'll frequently use the digital copy as a kind of preview, then purchase it once I see it's a good book I want to read.)

I've only been reflecting on this difference for a few minutes, but, to me, I think the major difference boils down to:

  1. Netflix series (movies, albums, etc) are non-essential, fictional works that take a long time to produce - think: fancy chocolates and caviar.
  2. News, generally, contains timely, important information - more meat and potatoes.
  3. While much of the super-critical news is not paywalled (e.g., product recalls, election dates, COVID stats, etc), a lot of information that is advantageous to know (discussions on interest rates, details on legislation, etc) is paywalled, compounding information asymmetries.
Sure, "stealing bad", but, IMO, someone stealing rice and beans from WalMart to feed their family is a different class of offense than someone robbing a boutique bakery because they can't get enough chocolate cake.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#512
post #363

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

> I think that's something worth reflecting on, about why we feel it's OK to pirate news articles, but not other IP As you noted it is not the norm to post pirate links here for IP other than news articles, but that doesn't mean that a lot of people think it is not OK to pirate those other forms of IP. In nearly any big discussion that even remotely involves video streaming there will be numerous posts from people ex…

Piracy is different from plagiarism.

People are understandably angsty about someone stealing credit. A NYT article is going to be a NYT article, not laundered around and presented as someone else's work.

Plus, there's the angle of enshitification and ads being injected into a paid service, and so on.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#513

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

To me, there is a sense that the news, which is real information about the society that we currently live in, should be availabe to all participants of that society. The notion of being a good citizen requires that one stays informed. Books, movies, videogames etc. don't have that role and are more consumption goods.

I'm of a similar mind. I take the more expansive view that everything created is part of our common property and that something like an LLM should be able to yield the summary and references to those creations. As I've said elsewhere, LLM systems might be our first practical example of an infinite number of monkeys typing and recreating Shakespeare (or the New York Times).

I understand that copyrights and patents are vehicles for ensuring a creator gets paid for their work, but they are flawed in not rewarding multiple parallel creations and that they last too long.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#514

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

> why we feel it's OK to pirate news articles, but not other IP

Who thinks this? I don't. I think copyright is wrong across the board. I would love if the same pattern of posting archive'd articles held for books, movies, et cetera.

I would love to change my mind on this, as it is a very unpopular opinion to have. But I have _never_ seen a morally or scientifically sound argument in favor of copyright law, and I've spent decades looking.

I think it subsidizes the creation of junk food content (superhero movies and clickbait news for example) while not contributing anything to the progress of science (paywalled scientific journals and textbooks). I shudder how much time I have wasted in my life consuming crap attention grabbing media and advertisements. I like to think if we lived in a world where everyone could be a publisher if they wanted to, the quality filters would be better, and information reaching us all would be more likely to be in our best interests.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#515

Earlier quoted context omitted.

> AI still could produce exact or extremely similar results of stuff it learned on. Can it do so more than a human can? I think that's the key here. If an AI is no more precise than a human telling you about the news article they read today then ChatGPT learning process probably can't be morally called copying.

So, if someone decompiles a program and compiles it again, it would look different. "It is not copying", we just did some data laundering. Feeding someone else data into your system is usually a violation of copyright. Even if you have a very "smart" system, trying to transform and obfuscate the original data.

> Feeding someone else data into your system is usually a violation of copyright

In some circumstances, yes, but often it's not, especially if you're not continuing to store and use it (which OpenAI isn't).

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#516

The NYT is preparing for a tsunami by building a sandcastle. Big picture, this suit won’t matter, for so many reasons. To enumerate a few: 1. Next gen LLMs will be trained exclusively on “synthetic”/public data. GPT-4V can easily whitewash its entire copyrighted training corpus to be unrecognizably distinct (say reworded by 40%, authors/sources stripped, etc). Ergo there will be no copyright material for GPT-5 to reg…

> rent seeking media companies Rent seeking? Media companies that actually create content are rent seeking? Versus the garbage hallucinations AI creates?

The New York Times is dying company that is rent seeking here. Along time ago, their content was valuable, yet now you can't even give it away to researchers.

I know because they tried to make a deal with my company, we passed because social media data is infinitely more valuable.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#517

Earlier quoted context omitted.

> AI still could produce exact or extremely similar results of stuff it learned on. Can it do so more than a human can? I think that's the key here. If an AI is no more precise than a human telling you about the news article they read today then ChatGPT learning process probably can't be morally called copying.

So, if someone decompiles a program and compiles it again, it would look different. "It is not copying", we just did some data laundering. Feeding someone else data into your system is usually a violation of copyright. Even if you have a very "smart" system, trying to transform and obfuscate the original data.

I'm regularly feeding other people's data into my "system" (brain) in order to produce my outputs.

So I'm a living breathing copyright violator. As a person I should be banned.

Fortunately, copyright is a bullshit fictitious right with no basis in natural law. So I don't lose much sleep over it.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#518
post #363

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

> I think that's something worth reflecting on, about why we feel it's OK to pirate news articles, but not other IP As you noted it is not the norm to post pirate links here for IP other than news articles, but that doesn't mean that a lot of people think it is not OK to pirate those other forms of IP. In nearly any big discussion that even remotely involves video streaming there will be numerous posts from people ex…

People used to leave newspapers in the trash, on the train, all over the place. Anyone could pick them up and read for free. I think it's reasonable for folks to carry this attitude into the digital age. People feel like news is something to share, it's not the source of creative expression, it's facts and as such we feel entitled to know the facts about our world and what is happening that might affect us.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#519
post #492

This, or a lawsuit like it is going to be the SCO vs IBM of the 2020's, to wit: a copyright troll trying to extract rent, with various special interests cheering it on to try and promote their own agenda (ironically it was Microsoft that played that role with SCO). It's funny how times have changed and at least now a louder group seem to be on the troll's side. I hope to see some better analysis on the frivolity of t…

>> There may be some commercial subtlety in specific cases that doesn't depend on scraping and training The key is to stop calling it "training" and use "learning" or just "reading". The argument from NYT will probably be that LLMs are just a fancy way to compress or abstract information and spit it back out. In which case "training" seems to support their case?

I don't recall the source, but when people read, they typically only remember 20% of what they read (or heard?). Machine training encodes much more than 20%, so it is much closer to copying than training. Now the emergent abilities that come from this could be considered learning and dare I say imagination (which is the opposite of copying).

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#520
post #190

Earlier quoted context omitted.

Isn't it totally normal to write articles / blog posts that effectively summarize, and often quote from, news articles?

My impression is that it’s not necessarily legal, but going after bloggers and proving damages based is just a huge waste of their time. OpenAI came by with their fat stack of funding and changed that.

No, it is very specifically and deliberately fair use. That is the primary intended purpose of fair use. The New York Times doesn't own the news; they just own their articles.
Post reply on HN