Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

591–600 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#591
post #52

Earlier quoted context omitted.

Sounds like you didn't read the article. Here's a better synoposis: I read a NYT article and publish an exact copy of that article on my website: copyright infringement. Train a model on NYT text and it outputs an exact copy of that text: also copyright infringement.

So presumably when they fix that issue (which, if the text matches exactly, should be trivially easy) then would you accept that as a sufficient remedy?

Basically, ya. It's not enough to change just a couple words around. But ya, there's probably some way to engineer around the problem.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#592
post #278

Earlier quoted context omitted.

Good comment, it was very funny to see how people desperately try to find moral justification for pirating media A but not B. "It's apples to oranges, you see, there are less letters in the NYT article than in the book and they are rendered differently, so it is ok to pirate their work. I did nothing wrong!" :)

I wonder what the reaction of some of the people who browse this forum would be if the output of their careers were so commonly pirated. Somehow, I think most think that this argument doesn't apply.

I’d be pretty delighted. I’m paid for getting projects done, not for keeping hold on some copyrighted code. I want all my code to be open sourced, and reused.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#593

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

How about if you read the paper every day and write opinion pieces about world events? Fair use?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#594
post #440

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

It would be nice to have a nice principled answer to this, but unfortunately, in our world, the answer is probably: if you start making LOTS of money doing this, they will come after you.

The best example is that sport scores, names and stats are not copyrightable by settled case law; however, you still have to go to the NBA and players union if you want to make a fantasy basketball game that has stats or names.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#595
post #50

Earlier quoted context omitted.

Read the article. It's not difficult to get ChatGPT to regurgitate recent, obviously copyrighted articles, verbatim.

It will be equally easy for ChatGPT to rewrite copyrighted content that makes the output materially different for a copyright claim to succeed also.

Then ChatGPT should do that.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#596
post #265

Earlier quoted context omitted.

To me, there is a sense that the news, which is real information about the society that we currently live in, should be availabe to all participants of that society. The notion of being a good citizen requires that one stays informed. Books, movies, videogames etc. don't have that role and are more consumption goods.

> should be available to all participants of that society. Who pays?

Every news source has biases. Under the paywall business model, the people who share the biases of their favored news outlets pay for them, and in exchange, they get to ensconce themselves inside a bubble free of dissenting viewpoints. This also reinforces the bias of the news outlet; if they don’t toe the line, they will lose subscribers.

Instead of paying news outlets to provide ourselves with filtered feeds of content that match our own biases, we could instead pay news outlets to produce competing streams of explicit propaganda to be freely disseminated. The overall bias and quality of the news would be largely unchanged, even if the biases were more obvious; in fact, it may even improve.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#597
post #363

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

> I think that's something worth reflecting on, about why we feel it's OK to pirate news articles, but not other IP As you noted it is not the norm to post pirate links here for IP other than news articles, but that doesn't mean that a lot of people think it is not OK to pirate those other forms of IP. In nearly any big discussion that even remotely involves video streaming there will be numerous posts from people ex…

> "subscribing is not an option because even though this paid service does exactly what I want now at a price that is trivial for me they might someday later change"

Gonna gamble and call bullshit on this.

My speculation: the most popular reason HN'ers give for pirating: they literally cannot get the content otherwise.

2nd most popular: it is such a pain to either to purchase the content or get it to run on bog standard software (like Firefox/Linux/etc.) that otherwise paying fans are driven to whatever the current equivalent is for bittorrent.

In fact, I don't believe I've ever seen a justification for using bittorrent or whatever due to what someone's favorite streaming service might do in the future. I'm assuming you saw at least one based on what you wrote-- care to give a link?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#598

Not that it would solve this, but how hard would it be for ChatGpt or other problems to cite the sources used in a response. Is that difficult to capture and tag to 'knowledge' within a LLM? It could be a best of both worlds type situation if LLMs cited sources and linked to the source itself. Isn't that what happened with Google News's home page? I seem to recall that when Google took it away in some markets, at the…

This is not possible. There is no database of sources inside an LLM. Just like the knowledge in your brain does not have sources attached.

For an example, you referenced "what happened with Google News's home page". Could you give me your source? You could probably search for some suitable article for a reference, but you don't know a source from your memory.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#599

Earlier quoted context omitted.

I'm of a similar mind. I take the more expansive view that everything created is part of our common property and that something like an LLM should be able to yield the summary and references to those creations. As I've said elsewhere, LLM systems might be our first practical example of an infinite number of monkeys typing and recreating Shakespeare (or the New York Times). I understand that copyrights and patents are…

a LLM is just a hugely lossy-compressed version of its training data, an abstraction of it. Much in the same way as when you read a book, your brain doesn't become a pirated copy of the text as you only store a hugely compressed version of it afterwards, a feeling for the plot, generated images and so on.

That's what I thought from my various readings about LLM systems. I'm guessing that the kerfuffle from the New York Times and other shortsighted organizations is that copyright allows them to control how their content is used. With humans, it's simple as its read and misremembered. Using it for LLM training requires a different model. It probably should be a RAND fee system based on volume of training data because, as you say, the training data is converted into an abstract form.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#600

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

What you described is entirely fair use, actually. Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost all of them are directly sourced from NYT and others. They'll say "so and so happened, according to The Times" (and usually link the article there)

> What you described is entirely fair use, actually

Just like during the pandemic how everyone became an epidemiologist, suddenly everyone's a copyright lawyer. I'll just dispute your assertion by saying:

1. Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most definitely not. A judge has wide latitude in determining fair use.

2. People should familiarize themselves with the four factors of fair use determination. In particular, if a work is purely derivative of a source work and substantially negatively impacts the market for the original work, it's very likely to not be considered fair use.

A great overview is https://fairuse.stanford.edu/overview/fair-use/four-factors/

Post reply on HN