Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

431–440 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#431

Earlier quoted context omitted.

Their ability to make money in the future is directly tied to their employers' ability to make money with their content. This is a closed financial loop. If OpenAI or any other AI company wants in, they should pay a licensing fee or get the laws changed, not just assume that they can take what they want and pretend like there are no negative consequences for the creator or the rights-holder.

No one is pretending there are no "there are no negative consequences for the creator or the rights-holder". Of course there are. But this is a story of rights-holders, who've already outgrown their usefulness, wanting to tap themselves into money stream they are not entitled to. ChatGPT isn't competing with NYT on a core competency . No one uses LLMs for original news reporting. They're obviously incapable of doing…

[deleted]

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#432
post #50

Earlier quoted context omitted.

For the paper or the author? What exactly was the licensing agreement for Op-Ed authors in 1962?

Read the article. It's not difficult to get ChatGPT to regurgitate recent, obviously copyrighted articles, verbatim.

It will be equally easy for ChatGPT to rewrite copyrighted content that makes the output materially different for a copyright claim to succeed also.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#433
post #363

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

> I think that's something worth reflecting on, about why we feel it's OK to pirate news articles, but not other IP As you noted it is not the norm to post pirate links here for IP other than news articles, but that doesn't mean that a lot of people think it is not OK to pirate those other forms of IP. In nearly any big discussion that even remotely involves video streaming there will be numerous posts from people ex…

Yeah, I don’t judge people for pirating or ad blocking, but the ludicrous justifications do get me - quite the entitled mental gymnastics. They remind me of bitcoin people trying to explain how mining is good for the environment.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#434

The suit demonstrates instances where ChatGTP / Bing Copilot copy from the NYT verbatim. I think it is hard to argue that such copying constitutes "fair use". However, OAI/MS should be able to fix this within the current paradigm: Just learn to recognize and punish plagiarism via RLHF. However, the suit goes far beyond claiming that such copying violates their copyright: "Unauthorized copying of Times Works without p…

Yeah, no - that proposal is no good. The correct solution is to have machine learning be more like human intelligence. You can't ask me to plagiarize a New York Times article. Not because of prompt rule violation but because I just can't. It's not how humans train (at least most).

You can't, but there are some people who can quickly memorize entire pages of written text.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#435
post #418

People who think the examples the lawsuit are “fair use” need to consider what that would mean. We’re basically going to let a few companies consolidate all the value on the Internet into their black boxes with basically no rules … that seems very dangerous to me. I hope a court establishes some rules of engagement here, even if it’s not this case.

Scraping is legal, and this seems like a transformative work to me.

Returning the full text of an article verbatim seems to me like the opposite of "transformative."

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#436

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

Largely because "news" aka facts is not and should not be copyrightable, so while the style, and exact format of the article may be copyrightable, the facts contained within are not. This makes a news story copyright murky in the eyes of wider society unlike a clearly 100% creative work like a TV Show or Movie. Further the news themselves self cannibalize, how many stories are just rewrites of stories from other outl…

why it is OK for the Washington Post to copy the NY times, but not ok for OpenAI or Archive.org?

If the Washington Post printed an article from the NY Times nearly verbatim and without attribution, it would not be OK and surely they would take legal action.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#437
post #418

People who think the examples the lawsuit are “fair use” need to consider what that would mean. We’re basically going to let a few companies consolidate all the value on the Internet into their black boxes with basically no rules … that seems very dangerous to me. I hope a court establishes some rules of engagement here, even if it’s not this case.

[dead]

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#438

Earlier quoted context omitted.

I think it is an RLHF problem and that you are right - this will blow up in the faces of the NYT. Specifically, the NYT examples all seem to be cases where they asked the AI to repeat their articles verbatim? So they ask it to violate copyright and because it's a helpful bot with a good memory, it does so. Solution: teach the model to refuse requests to repeat articles verbatim. It's easily capable of recognizing whe…

This is not how copyright works though. The verbatim quoting of articles is because when people brought up these questions initially the argument was that the NN doesn't really contain the training data or really just in an abstract, condensed way that does not constitute copying of the content. This demonstrates that no, the NN actually does contain the full articles, copied into the NN. Do you think any normal pers…

Search indexes contain exact copies of the pages they index, and that isn't a copyright violation.

> Why should we let OpenAI get away with this?

IP rights, like other private property rights, are a compromise between creators and consumers. What "should" be the case is essentially an argument about what balance creates the best overall outcomes. LLMs, for now, require large amounts of text to train, so the question is one of whether we want LLMs to exist or not. That's really a question for Congress and not the courts, but it'll be decided in the courts first.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#439
post #418

People who think the examples the lawsuit are “fair use” need to consider what that would mean. We’re basically going to let a few companies consolidate all the value on the Internet into their black boxes with basically no rules … that seems very dangerous to me. I hope a court establishes some rules of engagement here, even if it’s not this case.

a court has established this already

in japan, where they said anything goes for ai

so its best to not to lose a competitive edge with things that people openly publish on the internet, if you put it out there for everyone to see then expect other people to use it

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#440

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

It would be nice to have a nice principled answer to this, but unfortunately, in our world, the answer is probably: if you start making LOTS of money doing this, they will come after you.
Post reply on HN