Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

731–740 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#732

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

What you described is entirely fair use, actually. Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost all of them are directly sourced from NYT and others. They'll say "so and so happened, according to The Times" (and usually link the article there)

A lawyer starts a conclusion like this:

"It could be fair use if conditions a, b, and c are met. Condition a means..." ;)

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#733

Won't hold in court. GPT is a platform mainly providing answer to private individuals asking. Is like you ask a professor a question and he answered verbatim what copyrighted materials available (due to photographic memory) word for word back to you. Now if you take this answer and write a book or publish enmass on blogs for example, then you are the one should be sued by NYT. If GPT use the exact same wordings and p…

Professors and schools get into legal problems when professors pirate and/or otherwise distribute content they don't have licenses for.

[deleted]

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#734
post #670

Earlier quoted context omitted.

What you described is entirely fair use, actually. Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost all of them are directly sourced from NYT and others. They'll say "so and so happened, according to The Times" (and usually link the article there)

it's fair use if you don't make money from your project no?

No.

If that were true, I could take a band that I hate, copy all of their music note-for-note, then release an exact copy on the market and undercut them by selling their entire discography for $0.01

Fair Use requires one of several enumerated activities, including satire, education, journalism. You can’t just copy content and hope that it passes Fair Use.

Hire a lawyer if you are unsure. But at least read the Wikipedia article on the subject if you are going to talk about it.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#735
post #669

Earlier quoted context omitted.

I see the exact opposite - any open source model is going to become prohibitively expensive to train if quality data costs billions of dollars. We’re going to be left with the OpenAI’s and Google’s of the world as the only players in the space until someone solves synthetic data.

This feels like a 1996 "music is too expensive for kids so they HAVE to pirate it."

NYT is seeking billions of dollars - I’m not sure that’s a fair comparison.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#736
Huh, is this a big misunderstanding?

The copilot screenshot they gave in the ars-technica article as well as many of the screenshots in the NYT article seems like it's actually displaying correct behavior for browsing the web.

In these cases the system is more or less acting as a user agent (browser). AFAICT the NYT server actually gave that data to the user agent when it asked politely (200 OK, presumably). The user agent then displayed it to the user, which the user agent may do in any way it deems fit or appropriate.

There's only one or two cases where this has gone against the user or user agent, in very specific circumstances. The server can eg say 403 Forbidden whenever it likes, so if it returns a 200 OK, what's a user agent to do other than believe it at its word?

The only twist is that this user agent is now Imbued With AI (tm)(r)(c) . I don't think that really makes a difference here. If that's all this is, then it's more related to legal fights over certain ad-blockers or readability, which have similar functionality.

* https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20... , eg. page 45; I mean it says "Model: Web Browsing" at the top, and "Finished browsing" right on the page. That particular subsystem is now integrated, so the UI/UX is different now, but IIRC the link was in the pulldown?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#737
Related. Others?

NYT sues OpenAI, Microsoft over 'millions of articles' used to train ChatGPT - https://news.ycombinator.com/item?id=38784194 - Dec 2023 (80 comments)

The New York Times is suing OpenAI and Microsoft for copyright infringement - https://news.ycombinator.com/item?id=38781941 - Dec 2023 (837 comments)

The Times Sues OpenAI and Microsoft Over A.I.’s Use of Copyrighted Work - https://news.ycombinator.com/item?id=38781863 - Dec 2023 (11 comments)

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#738

Earlier quoted context omitted.

> they aren’t legally the same thing anyway. They should be.

Why? A human being isn’t infinitely scalable; they’re just different. It’s the same thing as going to a movie theatre to watch a movie vs. recording it with a camera.

A human churning butter, spinning cotton, or acting as a bank teller isn't infinitely scalable either. This is orthogonal to the point.

Times change. We're industrializing information creation and consumption (the latter is mostly here already), and we can't be stuck in the old copyright regime. It'll be useless in very short order.

All this road bump will do will give the giant megacorps time to ink deals, solidify their lead, and trounce open source. Twenty years on, the pace of content creation will be as rapid as thought itself and we'll kick ourselves for cementing their lead.

This is a transitional period between two wildly different worlds.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#739

Earlier quoted context omitted.

If it’s lossy compressed how come they have verbatim content from NYT in there that’s easy to recall? That’s what the lawsuit is about.

Many humans have photographic memories. Not common, but not unheard of for people to be able to memorize long portions of text verbatim. For example, the Wikipedia article https://en.wikipedia.org/wiki/List_of_people_claimed_to_poss... contains several examples of people who were able to look at pages and recite them back. That is actually a much stronger ability than GPT since GPT has presumably looked at them 100 t…

Yes, and a car is fast horse. Your argument does not tell us anything about whether or not GPT should be legal.

Laws are created by people (not by computers reasoning that all analogies must be true). And fairness is an important part of that process.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#740

Earlier quoted context omitted.

What you described is entirely fair use, actually. Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost all of them are directly sourced from NYT and others. They'll say "so and so happened, according to The Times" (and usually link the article there)

> What you described is entirely fair use, actually Just like during the pandemic how everyone became an epidemiologist, suddenly everyone's a copyright lawyer. I'll just dispute your assertion by saying: 1. Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most defini…

> if a work is purely derivative of a source work

CliffNotes, Wikipedia, etc. have huge quantities of summarized copyrighted work.

Post reply on HN