Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

451–460 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#451
post #435

Earlier quoted context omitted.

Scraping is legal, and this seems like a transformative work to me.

Returning the full text of an article verbatim seems to me like the opposite of "transformative."

That's not what "transformative" means for copyright.

It's more like, is the new work a distinct expression, e.g. satire or commentary, based on the original.

You can reproduce the original verbatim and still be transformative by adding an element of critique.

Example: https://www.dmca.com/articles/akilah-obviously-vs-sargon-of-...

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#453
post #408

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

This tendency at Hacker News are also much more of a threat to The New York Times than what Open AI is doing. Even the places like blogs/Reddit/social media submissions that summarize the article and post the relevant quotes. Unlike the summary of a movie, summarizing all of the relevant parts of a news article is extracting almost all the value from it, and giving it away for free. And the vast majority of people re…

Yes, and for nonfiction, it's also true that it usually depends on the original article for credibility. (If it were an anonymous poster making up a news story, most people wouldn't believe it.)

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#454

Earlier quoted context omitted.

Another factor to consider is that neural nets can function as lossy compression, which becomes extremely evident when using models that are overfit. Sometimes they're so overfit that the compression isn't even lossy, and the data is encoded verbatim in the NN.

Yes, but this then hits against learning/understanding and compression being fundamentally the same thing . I can't think of a better way to argue in favor of "it's fine if human does it, therefore it's fine if LLM does it", than from the "lossy compression" angle.

I can’t think of a better way to argue in favor of “LLMs are copyright laundering machines” than from the humanness angle.

Humans have rights, software tools don’t.

If you grant an LLM the full set of human rights, then it can consume information, regurgitate copyrighted works, and use it to generate money for itself. However, considering blatantly obvious theft as “homage” goes hand in hand with free will, agency, being in control of yourself, not being enslaved and abused, etc. Pondering various scenarios along those lines really gets to the heart of why an LLM is so very much not a human, and how subjecting it to the same treatment as humans is a ridiculous notion.

If you don’t grant LLM human rights, then ClosedAI’s stance is basically that pirating works is OK because they pass them through a black box of if conditions and it leads to results that they can monetize. That’s such a solid argument, it’ll surely play well in the court of law.

Training data is not an “LLM does it”; first because “it” here is not “learning” or understanding in human sense (otherwise you would have to presume that an LLM is a human), and second because a software tool doesn’t have agency and it’s really just Microsoft using a tool based on copyrighted works to generate profit.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#455
post #429

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

but what if they were also scraping, for example, Netflix content to use as part of their training set? There were some tweets the other day about how Midjourney could be prompted almost-exactly reproduce some frames of the film Dune. It wouldn't be shocking if these companies were using large databases of movies, with questionable legal status.

I see this a lot, and they very well may be. But, watch any behind the scenes documentary about any artsy movie and 9 out of 10, the director's will be waxing poetic about their inspirations, often include older movies or paintings which have uncannily similar scenes/frames. So it also wouldn't be shocking if a model trained on the same inspirations as the filmakers generates almost-exact frames as the movie makers.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#456

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

As someone pointed out, plenty of blogs made money off of doing just that. Many people go to Reddit to read news article summaries (and often a comment just pastes the whole article verbatim), instead of paying a site like the New York Times. Twitter and other social media sites are full of people summarizing articles from the New York Times. Any late breaking news article from Wikipedia is going to be mostly summarizing information from reporters.

I think people severely underestimate how much they've grown accustomed to this information being freely available. It's easy to say "Well it shouldn't be available with ChatGPT," but if we actually put everything back behind a paywall and stopped people from doing things like writing blogs or newsletters that summarize the news, people here would get angry very fast.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#457
post #363

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

> I think that's something worth reflecting on, about why we feel it's OK to pirate news articles, but not other IP As you noted it is not the norm to post pirate links here for IP other than news articles, but that doesn't mean that a lot of people think it is not OK to pirate those other forms of IP. In nearly any big discussion that even remotely involves video streaming there will be numerous posts from people ex…

> In nearly any big discussion that even remotely involves video streaming there will be numerous posts from people explaining why they pirate (usually with ridiculous justifications like "subscribing is not an option because even though this paid service does exactly what I want now at a price that is trivial for me they might someday later change").

I’ve read and participated in many such threads and I’ve literally never seen this take. Often what I see is complaints about having to learn different UI for different services/apps, no offline, ads injected into paid services, having to figure out which service a show is on, and generally terrible UI you can’t change/fix.

I don’t think I’ve ever really seen someone use the argument “yes it’s great today but they might charge more later”. Not saying people haven’t said that but it’s far from the main thing people say in my experience.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#458
post #190

Earlier quoted context omitted.

Isn't it totally normal to write articles / blog posts that effectively summarize, and often quote from, news articles?

My impression is that it’s not necessarily legal, but going after bloggers and proving damages based is just a huge waste of their time. OpenAI came by with their fat stack of funding and changed that.

No, in US law at least there can be no copyright of facts, only presentation. If you convey the same facts in different words that isn't a matter of fair use, it's never even a matter of copyright in the first place.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#459
post #435

Earlier quoted context omitted.

Returning the full text of an article verbatim seems to me like the opposite of "transformative."

That's not what "transformative" means for copyright. It's more like, is the new work a distinct expression, e.g. satire or commentary, based on the original. You can reproduce the original verbatim and still be transformative by adding an element of critique. Example: https://www.dmca.com/articles/akilah-obviously-vs-sargon-of-...

I don’t think the examples shown reflect an element of critique.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#460

The NYT is preparing for a tsunami by building a sandcastle. Big picture, this suit won’t matter, for so many reasons. To enumerate a few: 1. Next gen LLMs will be trained exclusively on “synthetic”/public data. GPT-4V can easily whitewash its entire copyrighted training corpus to be unrecognizably distinct (say reworded by 40%, authors/sources stripped, etc). Ergo there will be no copyright material for GPT-5 to reg…

If Microsoft doesn't get royalty free rights to resell access to everyone's content on demand, China will become the powerhouse of interference-free media? Rrrrrright....
Post reply on HN