Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

221–230 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#221
post #88

Earlier quoted context omitted.

Suit claims that GPT reproduced passages from NYT almost verbatim.

I'm sure the NYT uses dictionaries, encyclopaedias and style books verbatim as well. And they don't invent the facts they write about. As journalists they are compiling and passing along other knowledge. You usually don't get a piece of their income when a journalist quotes you verbatim (people usually don't get paid for interviews).

If the NYT reproduces other content verbatim too much, it will get in trouble.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#222

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

Using similar logic NYT should pay all actors involved in their articles.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#223
post #201

Earlier quoted context omitted.

That is an apples to oranges comparison. An article about a video/book would have the relevant information in text form without needing to show the video "here is the new stuff shown in Apples 2 hour long WWDC keynote". If not is common that a comment in the discussion gives a summary as a tl;dr With text articles behind paywalls the relevant information is hidden and only hinted at as a teaser.

To make it an apples to apples comparison, look at submissions where the link submitted is the retail link to the IP. For example, look at all the book link submissions on AMZN... https://news.ycombinator.com/from?site=amazon.com None of these have the Pirate Bay or Library Genesis or Anna's Archive or the equivalent as the top comment. Compare that to... https://news.ycombinator.com/from?site=nytimes.com And almost…

I wonder if this is because the purpose of linking to a book is to share awareness of that book’s existence - nobody is about to go and read it then and there to comment on its contents. Whereas the purpose of an article is to discuss it now, in the comments - the consumption horizon and bulk of the content is different.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#225

I read about this in the Times today (and am surprised that it wasn't on HN already). My guess is that the court will likely find in the Times favor, because the legal system won't be able to understand how training works and because people are "scared" of AI. To me, reading a book, putting it in some storage system, and then recalling it to form future thoughts is fair use. It's what we all do all the time, and I th…

It's extremely speculative to claim that LLM models are basically doing what humans do. There is very clearly something that isn't right about that because in order for a human to learn to speak and converse and they don't need to imbibe the entire corpus of all written text in human history - which is basically what we're doing with these LLMs. What we're giving them is vast amounts of data which is totally unlike how humans work. There's very clearly some gap here between what a LLM is doing and what a human is doing. So you can't use that as a basis to justify why it's ok for OpenAI to operate like this.

To put it another way, let's say I turn the dial all the way the other way, I train the worlds crappest LLM on NYT material, it massively massively overfits and all it will ever return is verbatim snippets of the NYT. Is that copyright infringement?

The core part of the argument here is actually just that OpenAI doesn't want to adhere to what the current standard is for using copyrighted material, if you want to use it and create something new with it you need to license the material. Since OpenAI's LLM isn't actually like a human it needs to license such a vast dataset that it would be uneconomical to run the business without stealing all the content.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#226

Earlier quoted context omitted.

> But also whatever the number 3 or 4 most valuable company in the world doesn’t get to scrape your content daily to repackage and sell as intelligent systems. Here's a thing though: for 99%+ of that content, being turned into feedstock for ML model training is about the only valuable thing that came of its existence . If it were not for world-ending danger of too smart an AI being developed too quickly, I'd vote for…

Except if you do that, you will see the number of content producers plummet quite quickly, and then you won't have any new training data to train new LLMs on.

Would it not logically follow that nothing of value would be lost, even if that were the case? From the point of view of LLMs and content creators, I would treat potential loss of future content being created like I would treat a lost sale. LLMs have value now because of training performed on content that already exists. There must be diminishing returns for certain types of content relative to others. Certain content is only of value if it is timely, and going forward, content that derives its worth from timeliness would find its creation and associated costs of production and acquisition self-justifying. If content isn’t of value to humans now or in the future, nor even of value to LLMs now or in the foreseeable future, not even hypothetically, then why should we decry or mourn its loss or absence or failure to be created or produced or sold?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#228
Fair use is something Wikipedians dance around a fair amount. It also meant I did a lot of reading about it.

It’s a four part test. Let’s examine it thusly:

1. Transformative. Is it? It spits out informative text and opinion. The only “transformation” is that its generative text. IMO that’s a fail.

2. Nature of the work - it’s being used commercially. Given it’s being trained partially on editorial, that’s creative enough that I think any judge would find it problematic. Fail on this criteria.

3. Amount. It looks like they trained the model on all of the NYT articles. Oops, definite fail.

4. Effect on the market. Almost certainly negative for the NYT.

IMO, OpenAI cannot successfully claim fair use.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#229

Earlier quoted context omitted.

I find “4nn4’$ 4rch1v3 dot ORG” actually way better than pirate bay for pirating knowledge. It’s amazing the amount of books that copyright laws prevent us from finding https://www.theatlantic.com/technology/archive/2012/03/the-m...

Sure. It's just curious to me that news article have a pirated knowledge link as the de facto top comment, but link submissions to, for example, books for sale on Amazon don't have a link to Anna's Archive or equivalent.

I think the archive of an article is more preservation of history and maintaining records of events which often disappear if not archived. The number of threads referencing articles which are defunct is always increasing. A book or movie or original content on the other hand will continue to hold its own commercial value so reproducing it is more akin to an actual loss for the license holder.

Definitely a grey area when that content is then used to train models though.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#230

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

If NYT was a HN startup the link to the archived version would be banned and dang would be slamming the ban hammer.

Please don't post baseless accusations. I think dang has said that he tries to moderate less, not more, when YC companies are involved. (Although it's impossible to say what he would do in this situation.)
Post reply on HN