Earlier quoted context omitted.
Suit claims that GPT reproduced passages from NYT almost verbatim.
I'm sure the NYT uses dictionaries, encyclopaedias and style books verbatim as well. And they don't invent the facts they write about. As journalists they are compiling and passing along other knowledge. You usually don't get a piece of their income when a journalist quotes you verbatim (people usually don't get paid for interviews).
NY Times copyright suit wants OpenAI to delete all GPT instances
221–230 of 921 posts
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#222If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#223Earlier quoted context omitted.
That is an apples to oranges comparison. An article about a video/book would have the relevant information in text form without needing to show the video "here is the new stuff shown in Apples 2 hour long WWDC keynote". If not is common that a comment in the discussion gives a summary as a tl;dr With text articles behind paywalls the relevant information is hidden and only hinted at as a teaser.
To make it an apples to apples comparison, look at submissions where the link submitted is the retail link to the IP. For example, look at all the book link submissions on AMZN... https://news.ycombinator.com/from?site=amazon.com None of these have the Pirate Bay or Library Genesis or Anna's Archive or the equivalent as the top comment. Compare that to... https://news.ycombinator.com/from?site=nytimes.com And almost…
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#224Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#225I read about this in the Times today (and am surprised that it wasn't on HN already). My guess is that the court will likely find in the Times favor, because the legal system won't be able to understand how training works and because people are "scared" of AI. To me, reading a book, putting it in some storage system, and then recalling it to form future thoughts is fair use. It's what we all do all the time, and I th…
To put it another way, let's say I turn the dial all the way the other way, I train the worlds crappest LLM on NYT material, it massively massively overfits and all it will ever return is verbatim snippets of the NYT. Is that copyright infringement?
The core part of the argument here is actually just that OpenAI doesn't want to adhere to what the current standard is for using copyrighted material, if you want to use it and create something new with it you need to license the material. Since OpenAI's LLM isn't actually like a human it needs to license such a vast dataset that it would be uneconomical to run the business without stealing all the content.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#226Earlier quoted context omitted.
> But also whatever the number 3 or 4 most valuable company in the world doesn’t get to scrape your content daily to repackage and sell as intelligent systems. Here's a thing though: for 99%+ of that content, being turned into feedstock for ML model training is about the only valuable thing that came of its existence . If it were not for world-ending danger of too smart an AI being developed too quickly, I'd vote for…
Except if you do that, you will see the number of content producers plummet quite quickly, and then you won't have any new training data to train new LLMs on.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#227Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#228It’s a four part test. Let’s examine it thusly:
1. Transformative. Is it? It spits out informative text and opinion. The only “transformation” is that its generative text. IMO that’s a fail.
2. Nature of the work - it’s being used commercially. Given it’s being trained partially on editorial, that’s creative enough that I think any judge would find it problematic. Fail on this criteria.
3. Amount. It looks like they trained the model on all of the NYT articles. Oops, definite fail.
4. Effect on the market. Almost certainly negative for the NYT.
IMO, OpenAI cannot successfully claim fair use.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#229Earlier quoted context omitted.
I find “4nn4’$ 4rch1v3 dot ORG” actually way better than pirate bay for pirating knowledge. It’s amazing the amount of books that copyright laws prevent us from finding https://www.theatlantic.com/technology/archive/2012/03/the-m...
Sure. It's just curious to me that news article have a pirated knowledge link as the de facto top comment, but link submissions to, for example, books for sale on Amazon don't have a link to Anna's Archive or equivalent.
Definitely a grey area when that content is then used to train models though.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#230It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
If NYT was a HN startup the link to the archived version would be banned and dang would be slamming the ban hammer.