Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

301–310 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#301

Earlier quoted context omitted.

Ruling in favor of copyright will call into question search engines and the like as well. Do you think Bing or Google are going to negotiate copying rights with the world's websites? LLMs are proving that intellectual property has a bunch of holes in it. It's been unstable ground to defend since day one. Upon what principle should we believe that one can own an idea and all performances or derivatives of it? Patents…

As previously said, search engines index and provide links. I’ll add that it constitutes fair use because a search engine isn’t itself a replacement for the articles that it indexes. But ChatGPT is actually providing an alternative that obviates the original articles themselves.

Search engines provide links, but also titles and snippets of the page -- enough for you to decide if you want to visit, and Google will show you their cached page if you ask for it.

Even the link is a copyrightable item -- artistic effort went into creating it

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#302

IANAL, but copyright protections are pretty much tied to content and format and not to the idea itself, with the intent of preventing (or putting a price on) the copying of original works. The Times will have a very hard time proving that their content is being re-marketed by OpenAI. Having a competing product based on your ideas. Compare: "Steve Jobs [was] a tyrant": https://www.nytimes.com/2011/10/07/technology/ste…

I think this is going to be a test of the fair use doctrine. https://www.copyright.gov/fair-use/ Now, there's this idea that "news" is just factual and therefore falls under "fair use". However, that's only part of what section 107 says. Fair use very much is still conditional, as there are 4 factors to be considered: (a) Purpose and character of the use, including whether the use is of a commercial nature or is for…

(a) looks bad, clearly commercial (b) ?? (c) looks bad, the LLM consumes "all" of the articles (d) looks bad, has a pretty significant impact on the market and value of the work

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#303

Earlier quoted context omitted.

Ruling in favor of copyright will call into question search engines and the like as well. Do you think Bing or Google are going to negotiate copying rights with the world's websites? LLMs are proving that intellectual property has a bunch of holes in it. It's been unstable ground to defend since day one. Upon what principle should we believe that one can own an idea and all performances or derivatives of it? Patents…

I think there is a major qualitative difference between generative AI and search engines. Search engines index the web and point you at other people's work, along the way showing perhaps too much of that content (thus "stealing" users from the target webpage). But they don't reshuffle existing content into something apparently new and original. The "malicious" case for generative AI is that it sucks in copyrighted wo…

... as opposed to a human doing the same thing ?

It's hard for me to reconcile that it's somehow OK for a student to write a term paper about something (e.g. "Ulysses"), Wikipedia doing the same thing, but on the other hand, not OK for chatGPT doing it.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#304
post #267

Earlier quoted context omitted.

> AI consuming copyrighted data and producing an output has to be considered a derivative work (or indeed, the model itself will be considered a derivative work) or IP protections are effectively broken. It’s not derivative work though. First, a human didn’t create it, so copyright protections don’t exist on its output. Machines don’t enjoy copyright protections, people do. It’s mechanically copying and reproducing p…

> First, a human didn’t create it, so copyright protections don’t exist on its output. Machines don’t enjoy copyright protections, people do. This is not correct. AI models are tools that humans use. This is like saying "it was typed on a computer therefore it doesn't enjoy copyright protections"

This has a ruling, see https://www.theartnewspaper.com/2023/05/04/us-copyright-offi...

When you're using ai models to generate the art, it's not considered human enough. If you then make a bunch of modifications to it, sure, but giving the initial prompt is currently insufficient

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#305
post #301

Earlier quoted context omitted.

As previously said, search engines index and provide links. I’ll add that it constitutes fair use because a search engine isn’t itself a replacement for the articles that it indexes. But ChatGPT is actually providing an alternative that obviates the original articles themselves.

Search engines provide links, but also titles and snippets of the page -- enough for you to decide if you want to visit, and Google will show you their cached page if you ask for it. Even the link is a copyrightable item -- artistic effort went into creating it

Search engines will also eventually stop serving the result if the source disappears. A LLM model that has been trained and published don't care at all about the source anymore.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#306

Earlier quoted context omitted.

But it also tells you everything on the page without needing to click It’s basically the “does this replace the original content” doctrine of fair use

Doesn't Google knowledge graph do that as well? Google is always giving me the answers I need before I click on a site. This was already normalized behavior prior to the existence of LLMs.

The knowledge graph is Wikipedia, that they use under license, or other sources that they have paid for a license to use.

What you're thinking of is "featured snippets". As far as I know, the justification behind those is that they are exact quotes that are followed by a citation (a link). Google argues those are fair use, since it's a properly referenced quote.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#307

Earlier quoted context omitted.

The part openai will have to argue is that it's not mererly compression but an irreversible transformation. Which is hard, best hope they have is trying to put the burden of proof on the nytimes to show you can make the model regurgitate their articles (with some nudging). If they manage that then nytimes is going to have a lot of trouble showing the model actually breaches their copyright, because just the informati…

Any form of lossy compression is an irreversible transformation. We do it all the time for video, audio and images (you can't recover the original data) and they are still copyrighted

when you compress a video, it doesn't recreate a new movie with a different story, different lines of text, different scenes and a different compositions for scenes that are similar to the "orginial".

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#308
post #272

Earlier quoted context omitted.

What's awful here? This way it's just a social interaction with both sides satisfied, where's the problem? I see a problem today, being spammed with shitty commercial "works" whereas I'd like to see something genuine and not just made for money.

If not for the ability to monetize, there would be a small fraction of the total available work out there. From music, to movies, to video games. And for many, the quality we come to enjoy just wouldn’t be possible. Do you think we’d have a Skyrim, or GTA, or equivalent if there weren’t millions to be made to employ thousands of people to make it happen? What about the largest and most influential films and TV shows…

You don't need to monetize via capital though -- you can pay people to do the creating directly. Eg. One of the biggest ways independent artists already get paid is via services like patreon

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#309
post #222

Earlier quoted context omitted.

I think this is going to be a test of the fair use doctrine. https://www.copyright.gov/fair-use/ Now, there's this idea that "news" is just factual and therefore falls under "fair use". However, that's only part of what section 107 says. Fair use very much is still conditional, as there are 4 factors to be considered: (a) Purpose and character of the use, including whether the use is of a commercial nature or is for…

Unfortunately we'll have 9 justices who know little about copyright law and nothing about tech to tell us what the law really is.

They will be significantly assisted by Supreme Court clerks, who are generally recent graduates (2-4 years) of the top law schools in the US. Your stereotypes of uninformed Congressmen from televised Congressional hearings don't really apply here.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#310
post #252

Earlier quoted context omitted.

> AI consuming copyrighted data and producing an output has to be considered a derivative work (or indeed, the model itself will be considered a derivative work) or IP protections are effectively broken. It’s not derivative work though. First, a human didn’t create it, so copyright protections don’t exist on its output. Machines don’t enjoy copyright protections, people do. It’s mechanically copying and reproducing p…

Seems unfair as we converge on AGI. If I memorize the lyrics to a song, is that a copyright violation? The lyrics are encoded in the arrangement of my neurons, after all.

Perhaps the difference is that you're a human doing it, and the other isn't?
Post reply on HN