Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

771–780 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#771
post #183

Earlier quoted context omitted.

> Just learn to recognize and punish plagiarism via RLHF OpenAI has created a $100bn company on this transfer. The Times may have an interest in a material fraction of that wealth.

The NYT is also worth a tiny fraction of that. If it looks like they might get anywhere, it might be better for OpenAI to buy them

If it looks like they might get anywhere, then lots of companies will also be able to get there, OpenAI can't buy all of them.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#772

Earlier quoted context omitted.

> What you described is entirely fair use, actually Just like during the pandemic how everyone became an epidemiologist, suddenly everyone's a copyright lawyer. I'll just dispute your assertion by saying: 1. Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most defini…

> if a work is purely derivative of a source work CliffNotes, Wikipedia, etc. have huge quantities of summarized copyrighted work.

First, you missed the "and". Do CliffNotes, Wikipedia, etc. substantially impact the market for the original work? For example CliffNotes does not - people who buy the CliffNotes version typically already have the original work as well (for example from coursework). And Wikipedia may well do more to interest people in the original work than to replace it.

Second, you ignored the "purely derivative" bit. You have to look at to what extent the use is derivative or transformative. See https://en.wikipedia.org/wiki/Transformative_use for a bit about that. (Note, this is a legal term defined by various precedents. OpenAI can't just argue, "Turning it into an LLM is a transform, so it is transformative!") Since CliffNotes is educational and Wikipedia is nonprofit, it is relatively easy for both to qualify as transformative.

As a result your response underscores the point that was made. There are a lot of shades of grey. You really can't just seize on a couple of phrases and key points, then jump straight to the answer. You have to understand how the courts will decide, and then accept that there is an actual judgment call whose outcome depends on the judge judging.

(I'm not a lawyer, but I have had excessive exposure to them in the past.)

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#773

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

> Something like, summarize all articles on US-UK relationships over past 5 years.

So like....Wikipedia, CliffNotes, encyclopedias, etc?

None of these pay royalties to original.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#774

The NYT is preparing for a tsunami by building a sandcastle. Big picture, this suit won’t matter, for so many reasons. To enumerate a few: 1. Next gen LLMs will be trained exclusively on “synthetic”/public data. GPT-4V can easily whitewash its entire copyrighted training corpus to be unrecognizably distinct (say reworded by 40%, authors/sources stripped, etc). Ergo there will be no copyright material for GPT-5 to reg…

About your 1. point: you can't possibly know that future models will be trained exclusively on synthetic data without any hit to performance. It is also not easy to reword the entire copyrighted training corpus without introducing errors or hallucinations. And you assume that this is just a fact?

Your second point reminds me a bit of 'War with the Newts' where humanity arms a race of sentient salamanders until they overthrow humanity. How could we not arm our newts if Germany might be arming theirs?

I also think basically everything else you wrote is wrong.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#775

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn :

> Implications: The Ninth Circuit's declaration that selectively banning potential competitors from accessing and using data that is publicly available can be considered unfair competition under California law may have large implication for antitrust law. [citation needed]

> Other countries with laws to prevent monopolistic practices or anti-trust laws may also see similar disputes and prospectively judgements hailing commercial use of publicly accessible information. While there is global precedence by virtue of large companies such as Thomson Reuters, Bloomberg or Google [or LexisNexis or Westlaw] effectively using web-scraping or crawling to aggregate information from disparate sources across the web, fundamentally the judgement by Ninth Circuit fortifies the lack of enforceability of browse-wrap agreements over conduct of trade using publicly available information.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#776

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

What you described is entirely fair use, actually. Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost all of them are directly sourced from NYT and others. They'll say "so and so happened, according to The Times" (and usually link the article there)

Sourcing, quoting, and linking is covered in the NYTimes content policy under fair use. See: https://help.nytimes.com/hc/en-us/articles/115014891408-Obta...

I think what wouldn't be covered is reproducing substantial portions of an article, especially if it's done without attribution. Tier 2 publications that fully reprint NYT or AP/Reuters articles are usually doing this via a paid News Service or Content License. See: https://nytlicensing.com/content/new-york-times-news-service...

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#777

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn : > Implications: The Ninth Circuit's declaration that selectively banning potential competitors from accessing and using data that is publicly available can be considered unfair competition under California law may have large implication for antitrust law. [citation needed] > Other countries with laws to prevent monopolistic practices or anti-trust laws may also see…

IANAL but aren't the key terms there "selectively banning" and "publicly available"?

NYT articles are largely behind a paywall for everyone. That means they are not publicly available, and a competitor who was blocked from accessing or reproducing that content without a license would not be "selectively banned"

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#778

Earlier quoted context omitted.

I would say 9 times out of 10 it's to get around the paywall and absolutely not some higher moralistic preservation of history. And everything is a grey area, determining the line is the existential purpose of these court cases. We've been here before with hyperlinking, then indexing and then linking with previews and the Canadian Facebook stuff but I think this has more standing.

If I buy a book, I get a work of literature. But if I buy a news subscription I get a series of facts riddled with advertisements. I accept the former, but I oppose the latter. I suspect I'm not the only one.

That's why you don't pay for news?

There are browser extensions that block ads. They are called ad blockers.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#779

Earlier quoted context omitted.

> What you described is entirely fair use, actually Just like during the pandemic how everyone became an epidemiologist, suddenly everyone's a copyright lawyer. I'll just dispute your assertion by saying: 1. Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most defini…

> suddenly everyone's a copyright lawyer Roll back 20+ years ago on Slashdot and you'll see the exact same thing. Copyright has been a hot button issue on the internet for decades. People end up thinking (rightly or wrongly) that they understand it without being a lawyer.

Legality aside I think copyright of digital things in the digital age is a net negative to humanity.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#780
post #756

Earlier quoted context omitted.

Are media really rent-seeking? They create new content and analysis, for which they want to be compensated. It seems quite different to hoarding natural resources or land, for example.

> It seems quite different to hoarding natural resources or land Indeed, it is quite different, because those things are scarce physical things in the real world. Intellectual property is a scam, and killing it once and for all will be one of the best things to come out of the current AI hype cycle. Nobody will "own" ideas, pieces of information, or strings of bytes.

Interesting. So as a hobby photographer I should only publicly release physical prints? An interesting idea.
Post reply on HN