Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

371–380 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#371
The NYT is preparing for a tsunami by building a sandcastle. Big picture, this suit won’t matter, for so many reasons. To enumerate a few:

1. Next gen LLMs will be trained exclusively on “synthetic”/public data. GPT-4V can easily whitewash its entire copyrighted training corpus to be unrecognizably distinct (say reworded by 40%, authors/sources stripped, etc). Ergo there will be no copyright material for GPT-5 to regurgitate.

2. Research/hosting/progress will proceed. The US cannot stop this, only choose to be left behind. The world will move on, with China gleefully watching as their biggest rival commits intellectual suicide all to appease rent seeking media companies.

3. Models can share weights, merge together, cooperate, ablate, evolve over many generations (releases), etc. Copyright law is woefully ill equipped to handle chasing down violators in this AI lineage soup, annealed with data of dubious/unknown provenance.

I could go on, but the point is that, for better or worse, we live in a new intellectual era. The NYT et al are coming along for the ride, whether they like it or not.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#372
post #84

Earlier quoted context omitted.

Making these things anathema to commercial interests and making training them at scale legally perilous would be a huge win.

A huge win for countries with lax copyright laws. These things aren't going away, the worst case scenario would be exactly that scenario playing out - then China (or some other peer to the US's tech sector) just continues developing them to achieve an economic advantage. All in addition to the obvious political implications of AI chatbots being controlled by them. The LLM genie is out of the bottle: an unfavorable co…

I don't give a shit about what China does.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#373

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

> Is that fair use?

As always, the answer is.. "it depends". I guess it depends mostly on the jurisdiction that applies to you. "Fair use" can have rather different legal meaning (or not exist at all) in different countries.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#374
post #213

Earlier quoted context omitted.

My impression is that it’s not necessarily legal, but going after bloggers and proving damages based is just a huge waste of their time. OpenAI came by with their fat stack of funding and changed that.

What parent poster meant is that it is normal that news organisations reference each other and report/cite/rephrase each other reports. For example all other news papers reported about the Watergate scandal reported by Bernstein&Woodward in the Washington Post.

Yeah but for every instance of that are face hugger links blogs that will rewrite the article and almost meant to deprive the source of any credit.

It’s not clear to me where the line is.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#375
post #84

Earlier quoted context omitted.

Making these things anathema to commercial interests and making training them at scale legally perilous would be a huge win.

> making training them at scale legally perilous would be a huge win. Why?

The open source people can continue to pretend they matter in this field and large corporations like Microsoft will stop stealing everything that moves on the internet.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#376
post #342

Earlier quoted context omitted.

> Just learn to recognize and punish plagiarism via RLHF. I'm not sure how your proposal would actually work. To recognize plagiarism during inference it needs to memorize harder. Kinda funny if it works though. We'd first train them to copy their training data verbatim , then train them not to. That is how it works, right? They're trained to copy their training data verbatim because that's the loss function. It's ju…

I wouldn't say it is an unexpected behavior. I remember reading papers about this memorization behavior few years ago (e.g., [1] is from 2019 and I believe it is not the first paper about this). It should be expected from OpenAI to know that LMs can exhibit memorizing behavior even after seeing the sample only once. [1] https://bair.berkeley.edu/blog/2019/08/13/memorization/

My expectation is that it can't memorize most of its training data. I expect it to memorize some.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#377

Earlier quoted context omitted.

They have content that LLMs want to use in training - millions of historical articles.

They created that content. It's an important distinction to make as compared to Reddit or Facebook where the users created the content.

The journalists created the content for the NYT, the users created it for Facebook. Both received something in return for their effort, and the content ended up being owned by NYT/facebook

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#378
post #84

Earlier quoted context omitted.

Making these things anathema to commercial interests and making training them at scale legally perilous would be a huge win.

>making training them at scale legally perilous Loading data to which you have no rights over into your software is legally perilous, yes. It's as easy as simply asking for and receiving permission from the data's rightsholders (which might require exchange of coin) to make it not legally perilous.

Sounds expensive.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#379
post #213

Earlier quoted context omitted.

What parent poster meant is that it is normal that news organisations reference each other and report/cite/rephrase each other reports. For example all other news papers reported about the Watergate scandal reported by Bernstein&Woodward in the Washington Post.

Those cite the original source that they used to write the article, the gpt models don't.

Depends on your prompt

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#380

The suit demonstrates instances where ChatGTP / Bing Copilot copy from the NYT verbatim. I think it is hard to argue that such copying constitutes "fair use". However, OAI/MS should be able to fix this within the current paradigm: Just learn to recognize and punish plagiarism via RLHF. However, the suit goes far beyond claiming that such copying violates their copyright: "Unauthorized copying of Times Works without p…

Yeah, no - that proposal is no good. The correct solution is to have machine learning be more like human intelligence. You can't ask me to plagiarize a New York Times article. Not because of prompt rule violation but because I just can't. It's not how humans train (at least most).
Post reply on HN