Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

581–590 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#581

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

Another factor to consider is that neural nets can function as lossy compression, which becomes extremely evident when using models that are overfit. Sometimes they're so overfit that the compression isn't even lossy, and the data is encoded verbatim in the NN.

> Sometimes they're so overfit that the compression isn't even lossy, and the data is encoded verbatim in the NN.

Here's an article from November 2023 that discusses this:

https://not-just-memorization.github.io/extracting-training-...

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#582
post #576

Earlier quoted context omitted.

That newspaper was likely paid for by someone, and could only be read by one person at a time.

And what if the person picking up the paper would stand up and shout the content of the article so all the people on the train would hear?

Reminds me of the movie News of the World. The main character's job is going from town to town, reading newspapers aloud.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#583
post #471

Earlier quoted context omitted.

Copyright doesn’t exist solely for the “promotion of useful sciences”. https://en.m.wikipedia.org/wiki/Copyright

Citing Wikipedia you already failed.. That is a General Article about Copyright world wide, I Specifically stated US Copyright, which is Authorized by Article I, Section 8, Clause 8 of the United States Constitution[1], implicitly for the promotion of the useful sciences. That is where congress derives its power to pass copyright laws, and to enforce copyright on the people of the United States. No other purpose is a…

You missed "and useful arts" in both of your comments. That's a key addition that you keep ommitting.

It is not just for sciences.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#584

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

> A sibling comment mentions search engines. I think there's a big difference. A search engine doesn't replace the source, not at all. Google has been accused for years of replacing sources with their "One Box"--the big answers at the top of the page, which are usually pulled from or corroborated by search results. They don't want you to leave the search results page (where the ads are).

Google is very careful to license all the content that shows up in that interface. They even pay Wikipedia, despite legally not needing to at all.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#585
Not that it would solve this, but how hard would it be for ChatGpt or other problems to cite the sources used in a response. Is that difficult to capture and tag to 'knowledge' within a LLM? It could be a best of both worlds type situation if LLMs cited sources and linked to the source itself. Isn't that what happened with Google News's home page? I seem to recall that when Google took it away in some markets, at the behest of the news orgs, they quickly reversed course as their traffic plummeted.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#586

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

If ChatGPT is based on neural networks, with no actual save-and-replicate facsimile behaviour, it no more "copies" original work than I do when I tell you about the news article I read today. I'd say the only real reason the Piratebay links thing you mentioned is not the norm is purely because those media sources have done a better job of striking fear into people doing that, so it's gone more underground. I.e. they'…

Copyright law does not care about the means of copying, just that you created something with substantial similarity to something you had access to. Whether or not the copy is in the form of a pixel array, blobs of random data being XORd to produce a full copy of music, or rows in a key/value attention matrix, doesn't matter.

Furthermore, there's Google research on extracting training set data from models. More specifically, Google found out that if you ask GPT to repeat the same word over and over again, forever, it eventually starts printing fully memorized training set data[0]. So it is memorizing stuff, even if it's not regurgitating it.

[0] When told of this, OpenAI's response was to block conversations with large amounts of repeated words in them.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#587

Earlier quoted context omitted.

Articles have ads on them, how are they not residual payments based on views?

I believe GP was referring to payments to the writer, not the publisher.

Yes, although I get that the route of the money may find it's way back to the journalist as salary. But generally goes into a pot for news gathering of which the salary will be withdrawn.

On ads it's acceptable to distribute them freely and it is advantageous to the company. Can we also see good journalism as an ad for the quality of a broader product?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#588

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

Because historically this is how news were shared. People would pick up a paper in a grocery store or cafe, read some of it, and leave it behind. They might rip out a page and take it home. Only one person paid and tens or hundreds gleam for free. This idea of sharing the story to nonsubcribers is as old as printed news itself. Instead news agencies prefer we forget that aspect of history, insist on being the “paper of record” while charging more money for easier to distribute media that gets sold globally. Yes, I think we are certainly not in the wrong here when we read the news for free.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#589

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

[deleted]

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#590

Not that it would solve this, but how hard would it be for ChatGpt or other problems to cite the sources used in a response. Is that difficult to capture and tag to 'knowledge' within a LLM? It could be a best of both worlds type situation if LLMs cited sources and linked to the source itself. Isn't that what happened with Google News's home page? I seem to recall that when Google took it away in some markets, at the…

not likely with the way these models have been trained - its basically broken down into sub-words that are all mashed together into probabilities.
Post reply on HN