Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

421–430 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#421
post #418

People who think the examples the lawsuit are “fair use” need to consider what that would mean. We’re basically going to let a few companies consolidate all the value on the Internet into their black boxes with basically no rules … that seems very dangerous to me. I hope a court establishes some rules of engagement here, even if it’s not this case.

Scraping is legal, and this seems like a transformative work to me.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#423
post #183

Earlier quoted context omitted.

> Just learn to recognize and punish plagiarism via RLHF OpenAI has created a $100bn company on this transfer. The Times may have an interest in a material fraction of that wealth.

The NYT is also worth a tiny fraction of that. If it looks like they might get anywhere, it might be better for OpenAI to buy them

OMG! Or they could just license the content. I suspect that would be both easier and less expensive. ;-)

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#424
post #281

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

The archive link doesn't threaten their jobs and helps them avoid paying for NYT. It's NIMBY, or rather it's true form of NIIIM (Not if it impacts me). Hypocrites are EVERYWHERE and are the majority.

It is pretty funny. If you go back and read the comments made yesterday about ChatGPT doing something much milder (using old articles to train data, some prompts fused to allow you to reproduce some of the articles though now don't work), you have a lot of comments talking about how The New York Times needs money and Open AI is using their work without paying for it.

Now a comment points out that HN News (and most of the internet) routinely does something much worse - allows people to bypass completely new articles in their entirety without paying - and almost all the comments are about how it's the New York Times fault for making it difficult to cancel subscription, the importance of news being available to everyone, the problems with copyright laws, etc.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#425
post #419

The NYT is preparing for a tsunami by building a sandcastle. Big picture, this suit won’t matter, for so many reasons. To enumerate a few: 1. Next gen LLMs will be trained exclusively on “synthetic”/public data. GPT-4V can easily whitewash its entire copyrighted training corpus to be unrecognizably distinct (say reworded by 40%, authors/sources stripped, etc). Ergo there will be no copyright material for GPT-5 to reg…

I'm sorry but this is such a bad take. Nice appeal to consequences. In my view, the New York Times is entirely justified in pursuing legal action. They invested time and effort in creating content, only to have it used without permission for monetary gain. A clear violation. Analyzing the factors involved for a "fair use" consideration: Purpose and Character of the Use: While the argument for transformation might hol…

Imo gpt itself is the transformative work.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#426

The NYT is preparing for a tsunami by building a sandcastle. Big picture, this suit won’t matter, for so many reasons. To enumerate a few: 1. Next gen LLMs will be trained exclusively on “synthetic”/public data. GPT-4V can easily whitewash its entire copyrighted training corpus to be unrecognizably distinct (say reworded by 40%, authors/sources stripped, etc). Ergo there will be no copyright material for GPT-5 to reg…

> rent seeking media companies

Rent seeking? Media companies that actually create content are rent seeking? Versus the garbage hallucinations AI creates?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#427
post #370

Earlier quoted context omitted.

To me, there is a sense that the news, which is real information about the society that we currently live in, should be availabe to all participants of that society. The notion of being a good citizen requires that one stays informed. Books, movies, videogames etc. don't have that role and are more consumption goods.

> which is real information about the society that we currently live in, should be availabe to all participants of that society. Who should pay the journalists or the investigative reporters?

The state, through taxes. It's a public good after all.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#428
post #366

Earlier quoted context omitted.

> , a movie, a video game, an album, a comic book, or any other form of IP, it is in fact very much _not_ the norm for the top-rated comment to be a Pirate Bay link. Probably because most print media is garbage and nobody in their right mind would actually pay to read them

> Probably because most print media is garbage and nobody in their right mind would actually pay to read them NYTs revenue keeps growing though.

Not from newspaper sales

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#429

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

but what if they were also scraping, for example, Netflix content to use as part of their training set?

There were some tweets the other day about how Midjourney could be prompted almost-exactly reproduce some frames of the film Dune. It wouldn't be shocking if these companies were using large databases of movies, with questionable legal status.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#430
> All of that costs money, and The Times earns that by limiting access to its reporting through a robust paywall.

Not to be pedantic, but NYT has the least robust paywall I've ever seen. Just turn on reader mode in your browser. Simple. I get that it's still tresspassing if I walk into an unlocked house, but NYT could try installing a lock that isn't made of confetti and uncooked pasta.

Post reply on HN