Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

351–360 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#351
post #265

Earlier quoted context omitted.

To me, there is a sense that the news, which is real information about the society that we currently live in, should be availabe to all participants of that society. The notion of being a good citizen requires that one stays informed. Books, movies, videogames etc. don't have that role and are more consumption goods.

> should be available to all participants of that society. Who pays?

Yes, someone needs to pay.

I see the gp post about pirating news as a very good point, while having no veleity to pay the New York Times, and being ok with not reading it in general.

But I also pay for my national (public) news outlet, and their articles are available to anyone anywhere in the world. I don't know how it should work, but I wish we could get to a system where the burden to keep news outlet alive is split thinly enough to have open but viable publications around the world.

Basically the same way weather stations collaborate all other the world and we pay for our local stations while getting acccess to all the forecast everywhere.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#352

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

> Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. [...] And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or any other form of IP, it is in fact very much _not_ the norm for the top-rated comment to be a Pirate Bay link.

If the story was linking directly to the "book, TV show, movie, video game, album, comic book, etc", and the link only worked for some people while others randomly got a login request or similar, you'd also see the top comment being a link to an archived version which avoids the login screen. That is: the main difference is that the archive link has the exact same content as the link submitted in the story, only bypassing the login screen that some people see. And the only reason the archive site has the content is that it didn't get the login screen; if everyone always got the login screen, what you would see on the archive site would be the same login screen.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#353

Earlier quoted context omitted.

That is a orthogonal to the discussion we were having. The topic was whether people should have free access to news, and how should it be financed, not the quality of that news. People have free access to public roads all around the world, and the quality wildly differs in that as well. Also the quality of for-profit news services does differ wildly, you might have an opinion about that of fox news, for example, but…

> That is a orthogonal to the discussion we were having. The topic was whether people should have free access to news, and how should it be financed, not the quality of that news. On the contrary, the quality of the news is very important to the discussion. There is no point in making trash freely available to the public, after all.

The topic is a bit more nuanced, and far wider than "not fitting my favourite narrative on some topics, so it is generally and objectively trash".

think about this: I will get mostly objective and useful reports of the flood approaching my home near the river regardless the narrative/interpretation they might have on some other topics, or the biased reporting on the merits of the government in handling the situation at the dams.

For me I'm not here to debate on the political policies of some governments, just gave a few examples of ways to fund public access to news. This discussion is over from my part.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#355

Earlier quoted context omitted.

I would be "happier" to pay a subscription to an aggregation platforms like hackernews or reddit to access archived articles that are linked to these sites. In turn a proportion of that could be passed on to the underlying publishers that I actually visit. I have nearly zero interest in reading articles that aren't linked to from an aggregation site. I don't want to read theguardian.com, or nytimes.com, or washington…

This is a common statement, but every attempt to sell that service has been a dismal failure. See for example blendle.

Nearly every attempt at starting a new aggregation site like hackernews or reddit has been a failure.

I’m not going to switch to a new website where no community exists just so I can pay for news articles. To work it needs to be integrated into an existing, successful aggregation website.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#356

Earlier quoted context omitted.

Yes, but this then hits against learning/understanding and compression being fundamentally the same thing . I can't think of a better way to argue in favor of "it's fine if human does it, therefore it's fine if LLM does it", than from the "lossy compression" angle.

Is there some LLM meta where understanding and compression are argued to be the same thing I’m not aware of? Anyone got more details on this? Superficially it sounds like total BS; a highly compressed zip file does not exhibit any characteristics of learning. Algorithmically derived highly compressed video streams do not exhibit characteristics of learning. ? I’ve vaguely heard the learning can be considered to exhib…

Even something as simple as LZW starts developing a dictionary. Not all compression is sufficient for understanding, but the more you compress a stream of data, the more dependent you are on understanding the source, because understanding the source allows you to take more shortcuts and still be able to reconstruct the data.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#357

Earlier quoted context omitted.

If ChatGPT is based on neural networks, with no actual save-and-replicate facsimile behaviour, it no more "copies" original work than I do when I tell you about the news article I read today. I'd say the only real reason the Piratebay links thing you mentioned is not the norm is purely because those media sources have done a better job of striking fear into people doing that, so it's gone more underground. I.e. they'…

So, if someone applies a filter to a video/audio, it is no more "copies" of the original work (no, it is still protected). AI still could produce exact or extremely similar results of stuff it learned on.

> AI still could produce exact or extremely similar results of stuff it learned on.

Can it do so more than a human can?

I think that's the key here. If an AI is no more precise than a human telling you about the news article they read today then ChatGPT learning process probably can't be morally called copying.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#358

Earlier quoted context omitted.

I agree with your general point but Hungary is probably the worst example you could have chosen from any EU country! The Orbán government is famously using it to spread propaganda and fake information in unprecedented levels. The level of control governments exert on public broadcasting networks is widely different. Since Meloni, the RAI in Italy is facing similar issues, but Hungary is still the canonic example of g…

That is a orthogonal to the discussion we were having. The topic was whether people should have free access to news, and how should it be financed, not the quality of that news. People have free access to public roads all around the world, and the quality wildly differs in that as well. Also the quality of for-profit news services does differ wildly, you might have an opinion about that of fox news, for example, but…

No it isn't an orthogonal discussion. The reason Orban wants people to have free access to his propaganda is because it directly serves his purpose. To finance it directly from sales of the media would defeat the purpose. Coupled with Orban's attack on free media it completes the picture.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#359

Earlier quoted context omitted.

I agree with your general point but Hungary is probably the worst example you could have chosen from any EU country! The Orbán government is famously using it to spread propaganda and fake information in unprecedented levels. The level of control governments exert on public broadcasting networks is widely different. Since Meloni, the RAI in Italy is facing similar issues, but Hungary is still the canonic example of g…

That is a orthogonal to the discussion we were having. The topic was whether people should have free access to news, and how should it be financed, not the quality of that news. People have free access to public roads all around the world, and the quality wildly differs in that as well. Also the quality of for-profit news services does differ wildly, you might have an opinion about that of fox news, for example, but…

I would argue the people of Hungary would be better off without hatred against asylum seekers and minorities, political opponents, lies and misinformation.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#360

NYT's perspective is going to look so stupid in future when we put LLMs into mechanical bodies with the ability to interact with the physical world, and to learn/update their weights live. It would make it completely illegal for such a robot to read/watch/listen to any copyrighted material; no watching TV, no reading library books, no browsing the internet, because in doing so it could memorise some copyrighted conte…

Memorising isn't the issue. It's providing it back verbatim and/or cutting access to the source.

You'd get the same problem with someone with a photographic memory who a group of people would turn to recite them the news instead of buying the newspaper.

As of now public performance of copyrighted material is infringement.

Post reply on HN