Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

781–790 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#781

Earlier quoted context omitted.

Copyright law doesn't mention opt outs or search engine snippet controls. It's not clear to me that robots.txt is the singular thing that makes Google legal. In US copyright law facts cannot be copyrighted, so copyright on factual content like newspaper articles is limited. Simply replacing a few words wouldn't work, but I am certain that GPT-4 is capable of paraphrasing factual content at a level that would not be c…

If I make a website that scrapes NYT and passes it back and forth through a machine translator, say, English -> Spanish -> English, then the content will be slightly modified. Is this legal to make money off of? Seems like the legal answer is unclear but, like Napster, such a system seems like it would lose in court.

It would be unlikely to be something you'd find paying customers for, though? I suppose if you charged a small percentage of what NYT charges people might be willing to consider it, but you'd have some costs for hosting etc., so I am skeptical about its viability as a business model...

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#782

The NYT is preparing for a tsunami by building a sandcastle. Big picture, this suit won’t matter, for so many reasons. To enumerate a few: 1. Next gen LLMs will be trained exclusively on “synthetic”/public data. GPT-4V can easily whitewash its entire copyrighted training corpus to be unrecognizably distinct (say reworded by 40%, authors/sources stripped, etc). Ergo there will be no copyright material for GPT-5 to reg…

Rule 1 of the Internet: If you put it on the Internet, it's not yours anymore.

You don't have to agree with it. You don't have to like it. But if you accept it and live by it, it's much harder to get burned.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#783
post #711

What if you were one of the people who read the Times from cover-to-cover every day and seriously tries to remember as much as possible because you consider it a trustworthy reference source? And if you were called upon to solve a problem based on knowledge you consider trustworthy, what would you come up with? What if you were even specifically directed to utilize only findings gleaned from the Times exclusively? An…

That would of course be fine. But then imagine that because human memory is not able to keep all that information straight, you made copies of all those newspapers. And then you started charging people for your knowledge. And then imagine that as part of your knowledge service, you would copy snippets from the times word for word and give that to your clients without citation and pass it off as your own.

Yup, that's the other side of the dodecahedron.

As I understand it, it's the copying that can lead to infringement.

Then again if you have acquired a legitimate copy, you should certainly be able to retain it and use it for reference.

But for training a model on someone else's data I wouldn't even want a copy.

Just skim the data and retain my own thoughts.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#784
post #749

Earlier quoted context omitted.

> What you described is entirely fair use, actually Just like during the pandemic how everyone became an epidemiologist, suddenly everyone's a copyright lawyer. I'll just dispute your assertion by saying: 1. Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most defini…

> Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most definitely not. A judge has wide latitude in determining fair use. You're the one presenting unfounded claims with confidence here. There is well established case law about not being able to copyright facts. If y…

> You're the one presenting unfounded claims with confidence here.

No, I'm not. On the contrary, I'm really looking forward to this case because I believe it will be a great test of a bunch of concepts that are totally novel in the world of copyright law as it applies to generative AI. The only things I am presenting with confidence are:

1. That anyone who declares that something is unambiguously fair use (or, contrarily, unambiguously infringing) is likely wrong. There is simply too much latitude by judges, and there have certainly been cases where a ruling went one way, only to be overturned on appeal.

2. While I certainly have an opinion on how I think this case will be decided, I'm not presenting that with unwarranted confidence. Instead, I linked that great article on the 4 factors of fair use determination because it's clear to me lots of people are saying "fair use!" on one side or the other with no understanding of the factors judges must actually consider when making a determination.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#785

The NYT is preparing for a tsunami by building a sandcastle. Big picture, this suit won’t matter, for so many reasons. To enumerate a few: 1. Next gen LLMs will be trained exclusively on “synthetic”/public data. GPT-4V can easily whitewash its entire copyrighted training corpus to be unrecognizably distinct (say reworded by 40%, authors/sources stripped, etc). Ergo there will be no copyright material for GPT-5 to reg…

Rule 1 of the Internet: If you put it on the Internet, it's not yours anymore. You don't have to agree with it. You don't have to like it. But if you accept it and live by it, it's much harder to get burned.

Rule 1 of the internet is "don't talk about /b/."

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#786

Earlier quoted context omitted.

What you described is entirely fair use, actually. Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost all of them are directly sourced from NYT and others. They'll say "so and so happened, according to The Times" (and usually link the article there)

> What you described is entirely fair use, actually Just like during the pandemic how everyone became an epidemiologist, suddenly everyone's a copyright lawyer. I'll just dispute your assertion by saying: 1. Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most defini…

I personally appreciate the semi truck sized loophole that is satire. One can include an entire copy written work within one's own work as long as the treatment of that other copy written work is parody / satire. This is a provision of US copyright law put in place to protect political satire, which can be anything, because politics is everything.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#787

Earlier quoted context omitted.

Yes, but this then hits against learning/understanding and compression being fundamentally the same thing . I can't think of a better way to argue in favor of "it's fine if human does it, therefore it's fine if LLM does it", than from the "lossy compression" angle.

I can’t think of a better way to argue in favor of “LLMs are copyright laundering machines” than from the humanness angle. Humans have rights, software tools don’t. If you grant an LLM the full set of human rights, then it can consume information, regurgitate copyrighted works, and use it to generate money for itself. However, considering blatantly obvious theft as “homage” goes hand in hand with free will, agency, b…

Humans don't exactly have the greatest track record of granting other humans rights. I don't presume they'll get it any better with AI.

What I expect to happen is whoever has the most influence and power will get what they want and we'll end up raising a generation with the implicit understanding of "that's just how things are," natural order, truth, reality, and all that jazz.

The only thing that ever changes outcomes is if the contradiction status quo is incapable of being managed.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#788

The suit demonstrates instances where ChatGTP / Bing Copilot copy from the NYT verbatim. I think it is hard to argue that such copying constitutes "fair use". However, OAI/MS should be able to fix this within the current paradigm: Just learn to recognize and punish plagiarism via RLHF. However, the suit goes far beyond claiming that such copying violates their copyright: "Unauthorized copying of Times Works without p…

Any lawsuit makes all the claims it can and demands every sort of relief it might plausibly have. That's not to say that's how it should be (it can have awful results), just to say that's what to expect (and hope courts only considers the reasonable claim - "stop freely sharing our data" and avoids ridiculous/anti-fair-use claim "you can't even store our data").

The thing about you claim, "Just learn to recognize and punish plagiarism via RLHF" is that we've had an endless series of prompt exploits as well as unprompted leakage and these demonstrate that an LLM just doesn't have fixed border between its training data and its output. This will it basically impossible for OpenAI to say "we can logically guarantee ChatGPT won't serve your data freely to anyone".

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#789
post #190

Earlier quoted context omitted.

Isn't it totally normal to write articles / blog posts that effectively summarize, and often quote from, news articles?

I think the issue is that they trained ChatGPT on the New York Times' proprietary IP without paying licensing fees and, the Times argues, that is illegal. By way of proof the Times has examples of ChatGPT dumping out articles verbatim.

This is exactly how I understand it. There’s a lot ink getting spilled about “summarizing isn’t illegal” and “what about Cliffs Notes” but that isn’t what this is about.

If the verbatim examples that have been going around are true, that’s bad. I’d love to know more details around it — prompts used, whether that’s an old model, etc. This seems like plagiarism more than anything.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#790

Earlier quoted context omitted.

They would still thrive but in other countries with other legal frameworks. The concept is way too valuable to disappear.

If its economically relevant us will use its iron fist to have its laws adopted the world over, like most things such as copyright or drugs

They could but that would pretty much mean giving up the tech supremacy to China since they won't apply it. China already doesn't care much about copyright so that's not going to stop them.

I suspect it wouldn't be too hard to convince the EU though, the EU has an history of giving up rights and markets to big copyright holders even if that hurts the local companies.

Post reply on HN