Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

251–260 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#251
post #183

Earlier quoted context omitted.

> Just learn to recognize and punish plagiarism via RLHF OpenAI has created a $100bn company on this transfer. The Times may have an interest in a material fraction of that wealth.

The NYT is also worth a tiny fraction of that. If it looks like they might get anywhere, it might be better for OpenAI to buy them

That would require NYT being willing to sell, which historically they have not been.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#252
post #120

Companies that have content all see dollar signs. NYT won't mind if you use their content to train LLMs - as long as they get a commission. Reddit will shut down their free API and make you pay to get training content. Discord is going to be selling content for AI training too - if they haven't already done so. Twitter is doing it. They didn't care before because LLMs were just experiments. Now we're talking trillion…

NYT do not "have" content, they create content. It's their raison d'etre.

They have content that LLMs want to use in training - millions of historical articles.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#253

We developers like to pretend that LLM's are akin to humans and that they've been using things like NYTimes like humans as educational material. But they are not. It's much simpler, proprietary writing is now integrated into the source code of OpenAI, it would be as if I would copy parts of other propriety code and copy paste it into my own codebase. Claiming copy paste is a natural evolving process of millions of ye…

> it would be as if I would copy parts of other propriety code and copy paste it into my own codebase. It's not copy-pasted; it's compressed in a lossy manner. Even GPT4 has nowhere near enough memory to store the entirety of its training data in a non-lossy compression format. Just likes how humans compress the information we read.

>Just likes how humans compress the information we read.

Humans don’t have the scale machines have and moreover humans aren’t sevices, that argument doesn’t fly.

I really think NYTs data isn’t that important and nor crucial, LLMs could’ve just elided it. However, it’s more about training on copyrighted data in general which is kind of crucial for OpenAi, they trained their LLMs indiscriminately on copyrighted content without any plan to share any profits.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#254

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

To me, there is a sense that the news, which is real information about the society that we currently live in, should be availabe to all participants of that society. The notion of being a good citizen requires that one stays informed. Books, movies, videogames etc. don't have that role and are more consumption goods.

I agree, but nothing worth having is free. NYT and other news outlets have to ultimately pay reporters to go out into the world and do the work. The reporters are not priests, and the NYT is not a church that lives off donations and tax exemptions. They need money to operate, and you may disagree with how they try to collect that money (paywall) but that doesn't solve their funding problem.

How would you pay for news otherwise?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#255

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

It is am ethical grey area, but if the paywall applied to all user agents, which would make it similar to say buying a Kindle book, then you might see that as pirating, whereas if you use an archive service that was served the HTTP response and cached it, then you are using a proxy UA.

If the news/magazine doesn't want this they can simple serve a cut down or zero length article to all non-paying viewers! But they want that SEO, and they want that marketing.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#257

Earlier quoted context omitted.

>or produced replicas from memory and sold them (the most comparable example), I’d be sued. This is not the most comparable example, because it's not what ChatGPT is doing. The most comparable example is if you were hired as a contractor and the employer asked you to write verbatim some copyright content you'd memorised. If the employer then published it, they'd be the one liable, not you. >Humans produce derivative…

> The most comparable example is if you were hired as a contractor and the employer asked you to write verbatim some copyright content you'd memorised. If the employer then published it, they'd be the one liable, not you. No, you'd both be liable. You are not allowed to create copies of a copyrighted work, even from memory, for any commercial purpose. Making it public or not is irrelevant. This is more obvious with s…

Your previous employer never bought AutoCAD, they licenced its use, paying a subscription. When you start working for them that licence was no longer available to you. So you would be unable to subsequently use it.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#258

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

If NYT was a HN startup the link to the archived version would be banned and dang would be slamming the ban hammer.

Like I said in another comment it is simpler than that. They just serve the login page/payment page to all HTTP requests. If they do that then the submission itself likely get's flagged as there is no workaround (just like if I submit my blog with a banner saying "hey you pay me $1 to read my cool post")

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#259

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

To me, there is a sense that the news, which is real information about the society that we currently live in, should be availabe to all participants of that society. The notion of being a good citizen requires that one stays informed. Books, movies, videogames etc. don't have that role and are more consumption goods.

Not everything is news that appears in a newspaper. There are opinion pieces, etc.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#260

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

It is am ethical grey area, but if the paywall applied to all user agents, which would make it similar to say buying a Kindle book, then you might see that as pirating, whereas if you use an archive service that was served the HTTP response and cached it, then you are using a proxy UA. If the news/magazine doesn't want this they can simple serve a cut down or zero length article to all non-paying viewers! But they wa…

Indeed there are media that are hard paywalled, e.g., the information. However these are prohibited on HN, which possibly create additional bias towards non-hard-paywalled publications
Post reply on HN