Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

871–880 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#871
post #771
post #183

Earlier quoted context omitted.

The NYT is also worth a tiny fraction of that. If it looks like they might get anywhere, it might be better for OpenAI to buy them

If it looks like they might get anywhere, then lots of companies will also be able to get there, OpenAI can't buy all of them.

They won't need to. Most don't have enough money to survive a prolonged round of lawsuits, and the potential damages are limited. The only real leverage is taking their models out of circulation and cutting their training set and that leverage only exist for the large publishers.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#872
post #658
post #623

Earlier quoted context omitted.

Adding an extra constraint of no copying verbatim from a very large and relevant corpus will be hard to guarantee without enormous databases of copyrighted content (which might not be legal to hold) and add an extra objective to a system with many often contradictory goals. I don’t think that’s the technology-sound solution or one in the interest of anyone involved. It’s much more relevant to license content from as…

What if OpenAI were to first summarize or transform the content before training on it? Then the LLM has never actually seen copyrighted content and couldn't produce an exact copy.

You are assuming a lossy compression. Stylistic guidelines and personal habits of beat journalists suggest you might not, depending on how detailed the LLM is. The complaint has many quotes that are long verbatim sections.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#873

Earlier quoted context omitted.

What if the LLM is running locally and doing all of these things rather than hosted on a webserver which is serving the content?

It doesn't matter, if everything else stays the same what matters is what it's used for. If it's used to make money, it would certainly hurt claims of fair use—maybe not for those that do the training, but for those that use it.

> If it's used to make money, it would certainly hurt claims of fair use

What if a human manually searches all those articles and transcribes / summarizes them to me in the way ChatGPT did?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#874
post #852

Earlier quoted context omitted.

I propose getting paid before doing the work for the actual labor of creation. Crowdfunding, patronage, comissions, sponsorships all seem like ethical ways to get things done sustainably. That way creators get paid before they work, not after. We must strengthen these business models that don't depend on artificial scarcity because this number selling nonsense was over the second computers were invented. It's as dumb…

How do you know what the value of the art will be before it's created? Guns N' Roses is a top 40 artist on Spotify nearly 35 years after producing an album. Should they not have been paid after 1991? If you argue that they were a popular band and therefore should have been paid accordingly up front, well what about their debut record, which sold 30 million copies? How would you predict that value before its creation…

You should be paid the accurate value of the labor. The pay should not scale more when no additional labor takes place.

This is how art worked for millenia; someone commissions a chapel roof painting, someone commissions a concerto, someone commissions a statue, someone buys a chair, etc.

Artists still do this today, and there is no issue determining value beforehand. Artists list their commission prices, or their hourly costs, etc. This is a perfectly normal thing that happens everyday.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#876

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

What you described is entirely fair use, actually. Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost all of them are directly sourced from NYT and others. They'll say "so and so happened, according to The Times" (and usually link the article there)

Correct. But those 2nd tier sources don't have the NYT copy verbatim. Do you really think the US NFL, as an example, would let OpenAI use all of its recorded games as a way to train some new GenAI game framework to build better American Football games? No. All that material is copyright. Public media is going to move to a very awkward era of ownership and licensing because all of these large companies looking to make a buck off public data sets are doing very little to make the economic model less one sided.

I hope the NYT prevails here, personally. Models will (and are) currently tainted by data they should not contain and for longer term privacy concerns this needs to be addressed early and have significant consequences or we're headed towards a world where this type of technology will make our ad-targeted world seem like a much more manageable past.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#877

Earlier quoted context omitted.

Fair use is specific to the US, as far as I'm aware. Moreover, Congress had to codify fair use (turn fair use common into statutory law in the form of 17 U.S. Code § 107) in order to make copyright statutes compatible with the First Amendment. Most other countries don't have freedom of expression and freedom of the press, so copyright law in a different country usually lacks a unifying exception test like fair use to…

> Most other countries don't have freedom of expression and freedom of the press This is demonstrably wrong. Many countries have both freedoms, albeit some have less strong protection than others.

Good point. I failed to qualify what I meant by freedom of expression, and made a meaningless claim regardless. Despite the US Constitution's relatively broad speech protections (e.g. don't criminalize hate speech, and allow truth as a defense to defamation claims), US governments don't always respect freedom of expression (e.g. KOSA would force social media companies to moderate more aggressively to "protect kids") or respect press freedom (e.g. police pepper spray journalists at protests). Even so, I think Congress wouldn't have bothered to codify fair use if the First Amendment weren't as broad as it is.

I replace the following sentence from my previous comment:

> Most other countries don't have freedom of expression and freedom of the press, so copyright law in a different country usually lacks a unifying exception test like fair use to supplement the specific enumerated exceptions.

with the following:

Copyright law in most countries usually lacks a unifying exception test like fair use to supplement the specific enumerated exceptions in each respective country.

The rest of my previous comment remains the same.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#878
post #799
post #749

Earlier quoted context omitted.

> Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most definitely not. A judge has wide latitude in determining fair use. You're the one presenting unfounded claims with confidence here. There is well established case law about not being able to copyright facts. If y…

The "unfounded claims" were backed up by a link to Stanford on fair use and copyright. That's the opposite of being unfounded. Remember. The NY Times does not have a record of filing frivolous lawsuits. Particularly not against companies with deep pockets. So it is almost certainly true that a lawyer who knows the law better than you thinks that this has a real chance. So you should be looking for flaws in trivial de…

The Stanford link is just generic information about the fair use tests and does nothing to backup the assertion.

> They aren't, in addition to facts they offer analysis, editorial positions, and so on.

Those opinions and ideas are also not copyrightable. Only expressions of them are copyrightable, which is why paraphrasing facts, ideas and opinions is not a violation of copyright.

> Fair use is filled with shades of grey.

Yes, but not all those shade are equal. There is a long history of litigation showing that paraphrasing news articles is fine.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#879
post #749

Earlier quoted context omitted.

> Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most definitely not. A judge has wide latitude in determining fair use. You're the one presenting unfounded claims with confidence here. There is well established case law about not being able to copyright facts. If y…

> You're the one presenting unfounded claims with confidence here. No, I'm not. On the contrary, I'm really looking forward to this case because I believe it will be a great test of a bunch of concepts that are totally novel in the world of copyright law as it applies to generative AI. The only things I am presenting with confidence are: 1. That anyone who declares that something is unambiguously fair use (or, contra…

You seem to be shifting the topic of this thread. The GP comment is about paraphrasing news articles while I don't see anything in the NYT lawsuit about paraphrasing. Rather, the NYT is concerned with exact reproduction or near exact reproduction. I too am very curious about the outcome of this case and wouldn't care bet either way on the outcome. I do have an opinion on what precedent would be better for our society but that doesn't mean I think that outcome is more likely.

However, none of that matters in this particular thread. There are well established precedents about paraphrasing news articles and they do not support the claim you made

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#880

Earlier quoted context omitted.

It doesn't matter, if everything else stays the same what matters is what it's used for. If it's used to make money, it would certainly hurt claims of fair use—maybe not for those that do the training, but for those that use it.

> If it's used to make money, it would certainly hurt claims of fair use What if a human manually searches all those articles and transcribes / summarizes them to me in the way ChatGPT did?

It might also be considered copyright violation, after evaluating the four fair use factors.
Post reply on HN