Earlier quoted context omitted.
Yes, but this then hits against learning/understanding and compression being fundamentally the same thing . I can't think of a better way to argue in favor of "it's fine if human does it, therefore it's fine if LLM does it", than from the "lossy compression" angle.
It's fine for a human to remember it. It's not fine for a human redistribute it for money (legally speaking). That's copyright infringement.
NY Times copyright suit wants OpenAI to delete all GPT instances
531–540 of 921 posts
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#532Earlier quoted context omitted.
If it can and does reproduce a piece of text verbatim then the text is indisputably stored somehow in the model.
That's just not true. There's no search and retrieval involved. It just associates the words so strongly in that context because they were in the training data so often that next-token prediction can (sometimes, in some limited circumstances) reproduce chunks of it. It's like if a human had read pieces of an article so many times and knew NYT style so well that they could spit out chunks of an article verbatim, but u…
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#533It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
I think the intent is really different. For LLMs you're essentially teaching them language by showing them lots of examples of written language - newspapers are of course a great example of written language. The goal of OpenAI is not to reproduce newspaper articles verbatim when asked questions (even if the answer could be a newspaper article) and the fact that it can happen is a side effect of how LLMs work. When a…
Seems like a nice split-the-baby resolution would be to send the NYT Corp a single article read amount anytime GPT plagiarizes more than what’s allowed at an academic institution.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#534It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
In 1990 it would have been considered normal and appropriate to clip an article out of a newspaper and post it on a communal corkboard. What are the key differences between that form of IP and others, and that analogy and the present situation of HN allowing archive links?
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#535Earlier quoted context omitted.
Yeah, no - that proposal is no good. The correct solution is to have machine learning be more like human intelligence. You can't ask me to plagiarize a New York Times article. Not because of prompt rule violation but because I just can't. It's not how humans train (at least most).
You can't, but there are some people who can quickly memorize entire pages of written text.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#536I love seeing all the AI sycophants squirm at this news. Here's to hoping NYT wins this one and gets everything they ask for, and more!
I don't use chat gpt to get the news but also i don't buy paywalls
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#537It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
> why we feel it's OK to pirate news articles, but not other IP Who thinks this? I don't. I think copyright is wrong across the board. I would love if the same pattern of posting archive'd articles held for books, movies, et cetera. I would love to change my mind on this, as it is a very unpopular opinion to have. But I have _never_ seen a morally or scientifically sound argument in favor of copyright law, and I've s…
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#538Earlier quoted context omitted.
> Well yeah, copying a work and using it for its original expressive purpose isn’t fair use, no? You have to use it for a transformative purpose. They transformed the weights. Just like reading the article transforms yours . As for verbatim reproduction, I'm pretty sure brains are capable of reproducing song lyrics, musical melodies, common symbols ("cool S"), and lots of other things verbatim too. Those quotes from…
This comment is just blatant anthropomorphizing of ML models. You have no idea if reading an article “transforms weights” in a human mind, and regardless, they aren’t legally the same thing anyway.
They should be.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#539It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
https://news.ycombinator.com/item?id=23735026
Even talking about it will get you scolded for talking about something "off topic"
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#540It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
Funny, I don't see it as a moral thing but more a "what can you get away with" thing. I fully assume that if I was to post a magnet link to a torrent for whatever the link was about, I would be banned. Morally speaking, I think it's perfectly reasonable to download a copy of something and either read the relevant info for my current task or to sample it to decide if I want to buy it. I see it no different to using th…
The problem with the thinking in the root comment is that it implicitly assumes that people’s behavior is morally consistent, or that they even try particularly hard to behave in a morally consistent way. That’s not really how people work. If you ask them to discuss morality in the abstract, they’ll try to come up with a consistent system. But their actual behavior is mostly dictated by social norms. And if you try to pin them down on the morality of their concrete actions, they’re more likely to stretch their moral system to accommodate their actions than the other way around.
None of this is to say anything about my own opinions on news sharing or OpenAI’s situation. It’s just that someone decrying piracy but also posting/sharing/upvoting links to copies of news articles is neither surprising, nor indicative of some deeper nuance to how people view morality around IP.