Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

331–340 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#331
post #230

Earlier quoted context omitted.

If NYT was a HN startup the link to the archived version would be banned and dang would be slamming the ban hammer.

Please don't post baseless accusations. I think dang has said that he tries to moderate less, not more, when YC companies are involved. (Although it's impossible to say what he would do in this situation.)

HN is currently facilitating piracy. Something your comment failed to address.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#332
post #229

Earlier quoted context omitted.

I think the archive of an article is more preservation of history and maintaining records of events which often disappear if not archived. The number of threads referencing articles which are defunct is always increasing. A book or movie or original content on the other hand will continue to hold its own commercial value so reproducing it is more akin to an actual loss for the license holder. Definitely a grey area w…

I would say 9 times out of 10 it's to get around the paywall and absolutely not some higher moralistic preservation of history. And everything is a grey area, determining the line is the existential purpose of these court cases. We've been here before with hyperlinking, then indexing and then linking with previews and the Canadian Facebook stuff but I think this has more standing.

If I buy a book, I get a work of literature. But if I buy a news subscription I get a series of facts riddled with advertisements. I accept the former, but I oppose the latter. I suspect I'm not the only one.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#333

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

I would be "happier" to pay a subscription to an aggregation platforms like hackernews or reddit to access archived articles that are linked to these sites. In turn a proportion of that could be passed on to the underlying publishers that I actually visit. I have nearly zero interest in reading articles that aren't linked to from an aggregation site.

I don't want to read theguardian.com, or nytimes.com, or washingtonpost.com, or bloomberg.com, I want to read news.ycombinator.com. Paying an individual subscription to every possible underlying site that could be linked to from news.ycombinator.com is a non-starter.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#334

Earlier quoted context omitted.

That's not true at all. If you pay someone to copy NYT articles for you verbatim, and then they give the copies to you, and then you publish them online, then you've both violated the copyright. You are never allowed to make copies of copyrighted works, even for private deals (making such copies for purely personal use, such as archival, falls under fair use - but you can't build a service out of that). So, if the su…

>But if you simply ask for a summary of a specific article by, say, just name and date, and you get a copy of it, it's clear that GPT is storing the original data in some way, and thus it has copied the NYT's protected works without permission. In this particular case they were using it via Bing, which actively did a HTTP request to the particular article to extract the content. So GPT hadn't memorised it verbatim, i…

The article states that they used it initially through ChatGPT, but that seems to have been fixed in the meantime, at least for the very simplistic queries that used to work ("the first paragraph of the Carl Zimmer article on old DNA" in ChatGPT used to return the exact data from NYT, and "next paragraph" could then be used to get the following ones). Even if this has been fixed, it still proves that ChatGPT encodes exact copies of NYT articles in its weights, which may be a violation in itself, even if it is prevented from returning them directly. Especially if they ever started distributing the trained model.

Additionally, even the use through Copilot is very debatable. They are not returning the NYT link, which requires a subscription, they are returning the contents of it even to non-subscribers. And they are doing this in a commercial product, not a non profit like the Internet Archive, which has some arguments for fair use.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#335

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

I would be "happier" to pay a subscription to an aggregation platforms like hackernews or reddit to access archived articles that are linked to these sites. In turn a proportion of that could be passed on to the underlying publishers that I actually visit. I have nearly zero interest in reading articles that aren't linked to from an aggregation site. I don't want to read theguardian.com, or nytimes.com, or washington…

This is a common statement, but every attempt to sell that service has been a dismal failure. See for example blendle.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#336

Earlier quoted context omitted.

Yes, but this then hits against learning/understanding and compression being fundamentally the same thing . I can't think of a better way to argue in favor of "it's fine if human does it, therefore it's fine if LLM does it", than from the "lossy compression" angle.

Is there some LLM meta where understanding and compression are argued to be the same thing I’m not aware of? Anyone got more details on this? Superficially it sounds like total BS; a highly compressed zip file does not exhibit any characteristics of learning. Algorithmically derived highly compressed video streams do not exhibit characteristics of learning. ? I’ve vaguely heard the learning can be considered to exhib…

Suppose you wanted to train an LLM to do addition.

An LLM has limited parameters. If an LLM had infinite parameters it could just memorize the results of every single addition question in existence and could not claim to have understood anything. Because it has finite parameters, if an LLM wants to get a lower loss on all addition questions, it needs to come up with a general algorithm to perform addition. Indeed, Neel Nanda trained a transformer to do addition mod 113 on relatively few examples, and it eventually learned some cursed Fourier transform mumbo jumbo to get 0 loss https://twitter.com/robertskmiles/status/1663534255249453056.

And the fact it has developed this "understanding" as an ability to learn a general pattern in the training data enables it to compress. I claim that the number of bits required to encode the general algorithm is fewer than the number of bits required to memorize every single example. If it weren't then the transformer would simply memorize every single example. But if it doesn't have space then it is forced to try to compress by developing a general model.

And the ability to compress enables you to construct a language model. Essentially, the more things compress, the higher the likelihood you assign them. Given a sequence of tokens say "the cat sat on the", we should expect "the cat sat on the mat" to compress into fewer bits than "the cat sat on the door". This is because the latter is far more common and intuitively more common sequences should compress more. You can then look at the number of bits used for every single choice of token following "the cat sat on the" and thus develop a probability distribution for the next token. The exact details of this I'm unclear on. https://www.hendrik-erz.de/post/why-gzip-just-beat-a-large-l... this gives a good summary.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#337

Earlier quoted context omitted.

That's not true. You can have Independent public broadcasting that is not owned by the government and is reporting critically on it.

It’s still a difficult tension. The government will always control the purse strings so independence is always going to come with conditions.

The Guardian in the UK is an example of an alternative: It is owned by a trust, which funds it.

Norway has substantial public media funding across the political spectrum, but as you point out it always comes with conditions, even is less so than the funding for the state owned broadcaster.

Combining the two models and putting public funds into several perpetual trusts intended to provide funding from their profits at arms length from any sitting government similar to the (private) trust funding The Guardian might be an interesting alternative.

(EDIT: Norway also has its own variation over The Guardian model - the second largest media group was founded by unions but is now majority owned by the combination of two public benefit trusts)

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#338

Earlier quoted context omitted.

I agree with your general point but Hungary is probably the worst example you could have chosen from any EU country! The Orbán government is famously using it to spread propaganda and fake information in unprecedented levels. The level of control governments exert on public broadcasting networks is widely different. Since Meloni, the RAI in Italy is facing similar issues, but Hungary is still the canonic example of g…

That is a orthogonal to the discussion we were having. The topic was whether people should have free access to news, and how should it be financed, not the quality of that news. People have free access to public roads all around the world, and the quality wildly differs in that as well. Also the quality of for-profit news services does differ wildly, you might have an opinion about that of fox news, for example, but…

> That is a orthogonal to the discussion we were having. The topic was whether people should have free access to news, and how should it be financed, not the quality of that news.

On the contrary, the quality of the news is very important to the discussion. There is no point in making trash freely available to the public, after all.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#339
post #330

Earlier quoted context omitted.

> , a movie, a video game, an album, a comic book, or any other form of IP, it is in fact very much _not_ the norm for the top-rated comment to be a Pirate Bay link. Probably because most print media is garbage and nobody in their right mind would actually pay to read them

I don't understand the downvotes - it's an extremely valid opinion. If people ask questions like that then they should be able to accept forthright answers? (It's the same reason for me. I have tried news site subs but eventually got so tired of the polemic that I cancelled. I won't sub again).

The obvious response is that if you don't like news and think it has no value then you don't have to read it.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#340

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

The real answer is it totally depends on whether your product grows to $10,000,000,000, and whether you pays part of it back. Search engines pay with referral traffic.
Post reply on HN