Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

531–540 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#531

Earlier quoted context omitted.

Yes, but this then hits against learning/understanding and compression being fundamentally the same thing . I can't think of a better way to argue in favor of "it's fine if human does it, therefore it's fine if LLM does it", than from the "lossy compression" angle.

It's fine for a human to remember it. It's not fine for a human redistribute it for money (legally speaking). That's copyright infringement.

Correct, just like it’s infringement to reproduce an article from memory using pen and paper intentionally. The person deciding to do that bears responsibility. OpenAI would be liable IFF they were intentionally facilitating that, instead of it being an undesired artifact from overfitting.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#532

Earlier quoted context omitted.

If it can and does reproduce a piece of text verbatim then the text is indisputably stored somehow in the model.

That's just not true. There's no search and retrieval involved. It just associates the words so strongly in that context because they were in the training data so often that next-token prediction can (sometimes, in some limited circumstances) reproduce chunks of it. It's like if a human had read pieces of an article so many times and knew NYT style so well that they could spit out chunks of an article verbatim, but u…

sort of like the idea of practice - repetition of something concentrates more brain space to that thing so the compression ratio of it can decrease and become less abstracted / more exact.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#533

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

I think the intent is really different. For LLMs you're essentially teaching them language by showing them lots of examples of written language - newspapers are of course a great example of written language. The goal of OpenAI is not to reproduce newspaper articles verbatim when asked questions (even if the answer could be a newspaper article) and the fact that it can happen is a side effect of how LLMs work. When a…

> The goal of OpenAI is not to reproduce newspaper articles verbatim when asked questions (even if the answer could be a newspaper article) and the fact that it can happen is a side effect of how LLMs work.

Seems like a nice split-the-baby resolution would be to send the NYT Corp a single article read amount anytime GPT plagiarizes more than what’s allowed at an academic institution.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#534

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

I would broaden the question beyond HN to society as a whole.

In 1990 it would have been considered normal and appropriate to clip an article out of a newspaper and post it on a communal corkboard. What are the key differences between that form of IP and others, and that analogy and the present situation of HN allowing archive links?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#535
post #434

Earlier quoted context omitted.

Yeah, no - that proposal is no good. The correct solution is to have machine learning be more like human intelligence. You can't ask me to plagiarize a New York Times article. Not because of prompt rule violation but because I just can't. It's not how humans train (at least most).

You can't, but there are some people who can quickly memorize entire pages of written text.

That's why I qualified with "at least most"

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#536

I love seeing all the AI sycophants squirm at this news. Here's to hoping NYT wins this one and gets everything they ask for, and more!

I don't know if winning this will improve their business model

I don't use chat gpt to get the news but also i don't buy paywalls

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#537
post #514

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

> why we feel it's OK to pirate news articles, but not other IP Who thinks this? I don't. I think copyright is wrong across the board. I would love if the same pattern of posting archive'd articles held for books, movies, et cetera. I would love to change my mind on this, as it is a very unpopular opinion to have. But I have _never_ seen a morally or scientifically sound argument in favor of copyright law, and I've s…

You can self-publish. Oh, you want to be able to publish other people’s work, and without their permission? How does that benefit the author?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#538
post #70

Earlier quoted context omitted.

> Well yeah, copying a work and using it for its original expressive purpose isn’t fair use, no? You have to use it for a transformative purpose. They transformed the weights. Just like reading the article transforms yours . As for verbatim reproduction, I'm pretty sure brains are capable of reproducing song lyrics, musical melodies, common symbols ("cool S"), and lots of other things verbatim too. Those quotes from…

This comment is just blatant anthropomorphizing of ML models. You have no idea if reading an article “transforms weights” in a human mind, and regardless, they aren’t legally the same thing anyway.

> they aren’t legally the same thing anyway.

They should be.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#539

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

It's not just tolerated, it's encouraged because "the alternatives suck worse"

https://news.ycombinator.com/item?id=23735026

Even talking about it will get you scolded for talking about something "off topic"

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#540

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

Funny, I don't see it as a moral thing but more a "what can you get away with" thing. I fully assume that if I was to post a magnet link to a torrent for whatever the link was about, I would be banned. Morally speaking, I think it's perfectly reasonable to download a copy of something and either read the relevant info for my current task or to sample it to decide if I want to buy it. I see it no different to using th…

I’d argue that morality always has a “what can you get away with” component. Things that are normalized tend to be seen as morally permissible, and things that are seen as abnormal are more likely to be seen as immoral.

The problem with the thinking in the root comment is that it implicitly assumes that people’s behavior is morally consistent, or that they even try particularly hard to behave in a morally consistent way. That’s not really how people work. If you ask them to discuss morality in the abstract, they’ll try to come up with a consistent system. But their actual behavior is mostly dictated by social norms. And if you try to pin them down on the morality of their concrete actions, they’re more likely to stretch their moral system to accommodate their actions than the other way around.

None of this is to say anything about my own opinions on news sharing or OpenAI’s situation. It’s just that someone decrying piracy but also posting/sharing/upvoting links to copies of news articles is neither surprising, nor indicative of some deeper nuance to how people view morality around IP.

Post reply on HN