Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

701–710 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#701
post #661
post #645

I don't think the lawsuit has any merit, but I'd still like to encourage Sam Altman et al, if they really care about the greater good, to go Keyser Söze and immediately release torrents of the weights and source code for GPT-4 under GPL.

AFAIK the IP deal with Microsoft only covers development before AGI . So at any point OpenAI could declare that a sufficient degree of AGI has been achieved and thus return to its philanthropic mission. With GPLed models and all. However, at this point the employees expect a multi-million cash-out for each of them. So the philanthropic mission seems to be gone out the window. And probably that’s also the way Sam Altm…

Luckily MS is now on the board so they’ll have a say in when AGI is declared

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#702

Earlier quoted context omitted.

What you described is entirely fair use, actually. Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost all of them are directly sourced from NYT and others. They'll say "so and so happened, according to The Times" (and usually link the article there)

>What you described is entirely fair use, actually. Based upon what? You think other publishers use NYTimes articles for free without license?

He's talking about citing and quoting NYTimes articles, not republishing them verbatim. That said, it's very different if you're a publication that sometimes cites reporting from other publications vs. a website exclusively dedicated to indexing and summarizing NYTimes articles.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#703

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

> If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? That's not a good question. If I look out of my window and see my neighbor go to the shop, that's fine. If I use cameras and track everybody I see on the street and put them in a database, then that's problematic and illegal in many places. Logic does not necessarily apply when scaling is involved…

It is a good question with a simple answer: no.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#704

Earlier quoted context omitted.

This comment is just blatant anthropomorphizing of ML models. You have no idea if reading an article “transforms weights” in a human mind, and regardless, they aren’t legally the same thing anyway.

> they aren’t legally the same thing anyway. They should be.

Why? A human being isn’t infinitely scalable; they’re just different. It’s the same thing as going to a movie theatre to watch a movie vs. recording it with a camera.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#705
post #645

I don't think the lawsuit has any merit, but I'd still like to encourage Sam Altman et al, if they really care about the greater good, to go Keyser Söze and immediately release torrents of the weights and source code for GPT-4 under GPL.

> I don't think the lawsuit has any merit The lawsuit fundamentally has merit. It asks a huge open question that no one knows the answer to. The outcome will be extraordinarily impactful. The question must be answered at some point. The case has merit even if NYT loses across the board.

For the good of the world, let's hope the NYT loses across the board. It's basically behaving like a copyright troll here.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#706

Earlier quoted context omitted.

> What you described is entirely fair use, actually Just like during the pandemic how everyone became an epidemiologist, suddenly everyone's a copyright lawyer. I'll just dispute your assertion by saying: 1. Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most defini…

> suddenly everyone's a copyright lawyer Roll back 20+ years ago on Slashdot and you'll see the exact same thing. Copyright has been a hot button issue on the internet for decades. People end up thinking (rightly or wrongly) that they understand it without being a lawyer.

One of my biggest gripes is a somewhat adjacent issue where everyone thinks they're an American copyright lawyer and that American copyright law is universal.

It's very possible that the example provided above is an example of fair use in some country, and that the website offering that service could be hosted there.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#707

Earlier quoted context omitted.

> OpenAI could declare that a sufficient degree of AGI has been achieved and thus return to its philanthropic mission The response from MSFT's legal team would be biblical if openai pulled this.

It’s literally in the contract such a distinction is at OpenAIs discretion

Technically firing CEO was also at the board's discretion, so I'm dubious whether that means anything at this point.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#708

Earlier quoted context omitted.

They're probably saying that because its what the supreme court said except about a human copying a work created by another human. https://www.npr.org/2023/05/18/1176881182/supreme-court-side...

That's a good bet. Down at the bottom of the linked PDF are some more interesting allegations: Count 5 - MS/OpenAI removed NYT copyright notices in violation of the DMCA. Count 7 - By attributing hallucinated garbage to NYT, MS/OpenAI is diluting NYT trademarks in violation of US Trademark law. I admit: I laughed. This will be an entertaining lawsuit to follow.

What will ultimately happen is that OpenAI and all big tech with have to pay out some sizable sum to large copyright holders, and in exchange be granted a de facto exclusive right to develop these technologies further because they’re the only ones who can do so “responsibly” with respect to copyright. It will take a long time to wind its way through the courts, but this could be the death knell for open source LLMs in the US.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#709

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

[flagged]

Can you please make your substantive points thoughtfully and without snark or putdowns?

Edit: it looks like you've unfortunately been breaking the site guidelines quite a bit lately. Can you please review them and stick to the intended use of the site? We'd appreciate it.

https://news.ycombinator.com/newsguidelines.html

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#710
post #703

Earlier quoted context omitted.

> If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? That's not a good question. If I look out of my window and see my neighbor go to the shop, that's fine. If I use cameras and track everybody I see on the street and put them in a database, then that's problematic and illegal in many places. Logic does not necessarily apply when scaling is involved…

It is a good question with a simple answer: no.

It depends. Google built a product out of scraping content (Google Search).

But what I'm saying is that answering the question does not allow you to deduce anything about your rights; that's what I mean by "not a good question".

Post reply on HN