Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

841–850 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#841
post #702

Earlier quoted context omitted.

>What you described is entirely fair use, actually. Based upon what? You think other publishers use NYTimes articles for free without license?

He's talking about citing and quoting NYTimes articles, not republishing them verbatim. That said, it's very different if you're a publication that sometimes cites reporting from other publications vs. a website exclusively dedicated to indexing and summarizing NYTimes articles.

I couldn't get gpt to quote an actual nyt article no matter how hard I tried...it just hallucinated in the general style of a news article.

Presumably, if it can remember at least a paragraph or two of each article, then surely the same would be true of any text it ingested and the model size would approach the dataset size (probably actually much larger). I don't believe this is the case at all, even searching around, I've not found any good recent examples of it regurgitating copyrighted text verbatim.

It's cool to hate AI stuff if you're a creative atm. But gotta love those generative/algorithm based PS brushes, that's still real art!

"Indeed, the opening paragraph of "A Game of Thrones" by George R.R. Martin, with the chapter titled "Bran," starts as follows:

"The morning had dawned clear and cold, with a crispness that hinted"

And then it cuts off, whether that's because OAI now have an oh shit filter or just the model had access to the first page or publicly available articles quoting the first line, I'm not sure.

I tried other chapters and random sections and it could get a sentence or two right but then hallucinated; what's more likely NYT and GRRM? That your works are being reproduced verbatim? Or that Facebook, YouTube descriptions, fan tumblrs and hell, the publicly available and multiple GoT related wikis that include a variety of passages from the books were used as training data?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#842
post #762

Earlier quoted context omitted.

> What you described is entirely fair use, actually Just like during the pandemic how everyone became an epidemiologist, suddenly everyone's a copyright lawyer. I'll just dispute your assertion by saying: 1. Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most defini…

>if a work is purely derivative of a source work This is the weakest part of the case(s) against OpenAI. "Derivative work" is a legal term of art meaning a direct adaptation, like writing a screenplay of a book or translating a book into another language. NYT has a stronger case than Sarah Silverman here because they can show actual 'memorized' text rather than just summarization, but given that those memorizations a…

"Transformative" seems to fit a lot more that "Derivative".

On the other hand, it's understandable why NYT is worried. OpenAI itself says that occupations like: Writers and Authors, Web and Digital Interface Designers, News Analysts, Reporters, and Journalists, Proofreaders and Copy Markers are "90-100% exposed" to what OpenAI is building.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#843

Earlier quoted context omitted.

> To keep things simple, let's say I never regurgitate chunks of verbatim NYT articles, maybe quite short snippets. You just described Google. When you think about it, it's surprising that Google is legal. However, it is well established that what Google does is perfectly legal. Remember that internally Google keeps and uses complete verbatim copies of every web page they index. Yes, Google offers a link to the sourc…

The reason why Google keeping entire digital copies of other people's copyrighted works is legal is because copyright is all about distribution rights. Any person can possess the entire works of Disney (without paying for them), for example and as long as they do not distribute those works they're 100% in the clear. Possession is not a crime when it comes to copyright. It's not like physical things (e.g. drugs or gun…

>Any person can possess the entire works of Disney (without paying for them)

Man thinks piracy is legal

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#844
post #418

People who think the examples the lawsuit are “fair use” need to consider what that would mean. We’re basically going to let a few companies consolidate all the value on the Internet into their black boxes with basically no rules … that seems very dangerous to me. I hope a court establishes some rules of engagement here, even if it’s not this case.

The opposite is also concerning. IP law has always been convoluted, messy, contradictory, and morally ambiguous. The complaints of IP violation by LLMs are simply taking these inherent flaws and making them immediately obvious, forcing decisions that ultimately will set precedents on the legality of human thought that I don’t think anyone will be comfortable with. People understandably see OpenAI and Microsoft as potentially dangerous to be given so much leeway, but fail to consider on the flip side companies like Disney who have already more or less dictated the majority of copyright law for decades now. They must be chomping at the bit at the legal precedents potentially coming down the pipeline that call into question the ability to interact with any kind of media or information at any level without potentially being on the hook monetarily.

I think all this is doing is making us realize that we have built a massive economic system on a fundamentally flawed idea of ownership over ideas, and the only two solutions will be to tear up the rule book, which will be extremely painful, or double down, which will be fatal.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#845

Earlier quoted context omitted.

> To keep things simple, let's say I never regurgitate chunks of verbatim NYT articles, maybe quite short snippets. You just described Google. When you think about it, it's surprising that Google is legal. However, it is well established that what Google does is perfectly legal. Remember that internally Google keeps and uses complete verbatim copies of every web page they index. Yes, Google offers a link to the sourc…

You took that quote out of context and missed the broader point in the process. The snippets provided in regular search results cannot generally replace the substance of the full articles they link to, while that's the whole point of GP's hypothetical website—it simply doesn't reproduce large chunks of text verbatim, presumably to avoid copyright infringement claims in the hypothetical's frame, and in GP's rhetorical…

>content whose substance was created by someone else

And how did the training data contribute to the content in any meaningful way? Inspiration isn't substance.

You think all fantasy writers gotta pay Tolkien estate bc so much of fantasy draws from his tropes? Lmao no.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#846
post #762

Earlier quoted context omitted.

>if a work is purely derivative of a source work This is the weakest part of the case(s) against OpenAI. "Derivative work" is a legal term of art meaning a direct adaptation, like writing a screenplay of a book or translating a book into another language. NYT has a stronger case than Sarah Silverman here because they can show actual 'memorized' text rather than just summarization, but given that those memorizations a…

"Transformative" seems to fit a lot more that "Derivative". On the other hand, it's understandable why NYT is worried. OpenAI itself says that occupations like: Writers and Authors, Web and Digital Interface Designers, News Analysts, Reporters, and Journalists, Proofreaders and Copy Markers are "90-100% exposed" to what OpenAI is building.

We should all be worried about that. If journalism is replaced with AI, truth is replaced with the AI hallucination du jour.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#847
post #194
post #72

Would be funny if NT Times won this and all commercial LLMs were shut down. Then LLMs would be distributed only via torrents, like most copyright infringing media.

What will happen in this case is that large content providers will get paid directly and smaller content providers will get rolled up into a licensing bag and get small indirect payouts. For example, we might see a model where people who's books have been used will get a pay out proportionate to the sales of the book (perhaps), so if your books sells just a few thousand copies expect $20 but if you sell millions expe…

What will happen is that all will go to China and maaaybe some third world country, or run your own models from shady sources.

So you will use a Chinese AI that spies on you, or you will use some shady service from a shady country (that will play cat and mouse like torrent sites).. or most likely you will run your own model when you are computer literate and no model if you are not.

Actually most models are so lobotomized allready that probably better to run your own, as long as you have a good enough computer.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#848

Earlier quoted context omitted.

> What you described is entirely fair use, actually Just like during the pandemic how everyone became an epidemiologist, suddenly everyone's a copyright lawyer. I'll just dispute your assertion by saying: 1. Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most defini…

> suddenly everyone's a copyright lawyer Roll back 20+ years ago on Slashdot and you'll see the exact same thing. Copyright has been a hot button issue on the internet for decades. People end up thinking (rightly or wrongly) that they understand it without being a lawyer.

This seems a bit disingenuous. Lawyers DISAGREE on this stuff (as we will see in this case) and a court will decide the reality by fiat.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#849

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

If there is payment then usually there is an agreement. An agreement can limit fair use. Can the NYT, via an agreement, e.g., "Terms of Use", limit what the subscriber does with the articles. There is not much precedent that suggests otherwise.

Consider the analogy from libraries that want to do data mining.

"Unfortunately, in licenses for digital scholarly content the majority of content acquired by research libraries publishers often include terms that prohibit certain uses that would otherwise be allowable under the Copyright Act. For instance, licenses may require libraries or individual researchers to negotiate for otherwise lawful activities, such as text and data mining, and to pay exorbitant fees on top of the cost of the content itself. While new regulations allow researchers to circumvent technological protection measures to access copyrighted materials, licenses for that content may include terms that explicitly prohibit this circumvention. In many cases, these activities might actually increase the value of published material; for instance, if a data-mining project yields new knowledge about a topic covered in a journal, it may very well spark new interest in that journals content. Libraries and publishers have often assumed that license terms that restrict copyright exceptions are enforceable under state contract law. There is, however, surprisingly little case law on this point."

https://www.arl.org/wp-content/uploads/2022/07/Copyright-and...

Putting some string in a robots.txt to try to stop data collection is an amusing "solution". Should copyright owners have "Terms of Use" that limit usage for commercial "AI" purposes.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#850

Earlier quoted context omitted.

For books, if it's a client reader software frustration, then you should still buy the digital version and then you can pirate the PDF book and use as desired within the constraints of copyright law (e.g. don't go sharing the PDF). That way you get the client you want but you still paid the content creator. But to use the argument, "oh, I don't like their client so I'm going to not pay them" is BS. For UFC, your comp…

The problem with buying by the crappy DRM version is that it provides no incentive to the publisher to change. I have thought about this long and hard, but ultimately the only way Spotify came about was because nobody bought the terrible DRM’d music the labels wanted to foist on us. We need to inflict the same pain for books. Personally, I think it would be preferable to donate the same amount to the Books Trust or y…

> The problem with buying by the crappy DRM version is that it provides no incentive to the publisher to change.

Then don't consume it and don't buy it. If you stop paying the abusive publisher, they'll be forced to change their policies.

The fact that you don't want to fund what is admittedly a rather abusive industry does not magically make it right to consume other peoples' work for free. That's theft-adjacent. You're not entitled to any piece of entertainment without paying for it.

Post reply on HN