Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

341–350 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#341

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

Good observation. I now wanna start commenting with pirate links to other media, but HN would tear me to shreds real quick I guess.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#342

The suit demonstrates instances where ChatGTP / Bing Copilot copy from the NYT verbatim. I think it is hard to argue that such copying constitutes "fair use". However, OAI/MS should be able to fix this within the current paradigm: Just learn to recognize and punish plagiarism via RLHF. However, the suit goes far beyond claiming that such copying violates their copyright: "Unauthorized copying of Times Works without p…

> Just learn to recognize and punish plagiarism via RLHF. I'm not sure how your proposal would actually work. To recognize plagiarism during inference it needs to memorize harder. Kinda funny if it works though. We'd first train them to copy their training data verbatim , then train them not to. That is how it works, right? They're trained to copy their training data verbatim because that's the loss function. It's ju…

I wouldn't say it is an unexpected behavior. I remember reading papers about this memorization behavior few years ago (e.g., [1] is from 2019 and I believe it is not the first paper about this). It should be expected from OpenAI to know that LMs can exhibit memorizing behavior even after seeing the sample only once.

[1] https://bair.berkeley.edu/blog/2019/08/13/memorization/

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#343

Earlier quoted context omitted.

Possibly because once an article is published the author receives no further payment. In all other mediums, there are residuals and royalties to be paid to the creators of the work.

Articles have ads on them, how are they not residual payments based on views?

I believe GP was referring to payments to the writer, not the publisher.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#344

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

it's similar to how easy it's to subscribe NY times and then how hard it's to unsubs. They require extra steps and it's well known. So They get what they deserve? Do you see the poínt. They are lie spreaders, nothing else

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#345

Earlier quoted context omitted.

I would be "happier" to pay a subscription to an aggregation platforms like hackernews or reddit to access archived articles that are linked to these sites. In turn a proportion of that could be passed on to the underlying publishers that I actually visit. I have nearly zero interest in reading articles that aren't linked to from an aggregation site. I don't want to read theguardian.com, or nytimes.com, or washington…

This is a common statement, but every attempt to sell that service has been a dismal failure. See for example blendle.

[deleted]

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#346

Earlier quoted context omitted.

It is actually pirating content by companies for humongous profit, or pirating by individual human beings for free access to culture and entertainment, oftentimes for content one has already paid for, but rendered inaccessible by megacorporations.

Which content making businesses earn humorous profit margins? Are all the journalist layoffs a fever dream? This is one of the more profitable ones, and only because they employ unscrupulous tactics: https://www.macrotrends.net/stocks/charts/NWS/news/profit-ma... This is NYT, the most successful news business: https://www.macrotrends.net/stocks/charts/NYT/new-york-times... As for movies/tv show/music makers, let’s ju…

The movie/tv show and music business can keel over and die tomorrow - it wouldn’t affect the value of art produced by humans at all. I see those more as exploitative leeches than as contributing anything positive.

If only piracy would actually harm these businesses but alas as often demonstrated it has zero effect on their bottom line, if anything it increases their profits.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#347
post #278

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

Good comment, it was very funny to see how people desperately try to find moral justification for pirating media A but not B. "It's apples to oranges, you see, there are less letters in the NYT article than in the book and they are rendered differently, so it is ok to pirate their work. I did nothing wrong!" :)

Of course pirating any media is totally fine from a moral standpoint.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#348
Let's try the "reverse the gender" card.

Let's say OpenAI was trained on all the Windows source code (without approval from MS).

GPT could pretty much replicate the windows code with even not that clever prompt by any user. "Write an OS CreateProcess function like Windows 10 source code would have."

It would infuriate MS to put it mildly, enough to start a lawsuit.

I know the license to the MS source code and NYT articles aren't the same.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#349

Earlier quoted context omitted.

Another factor to consider is that neural nets can function as lossy compression, which becomes extremely evident when using models that are overfit. Sometimes they're so overfit that the compression isn't even lossy, and the data is encoded verbatim in the NN.

Yes, but this then hits against learning/understanding and compression being fundamentally the same thing . I can't think of a better way to argue in favor of "it's fine if human does it, therefore it's fine if LLM does it", than from the "lossy compression" angle.

Humans are defined not just by their abilities but by their limitations too. We celebrate our achievements because sometimes they surpass the limitations of an average human.

Our collective human limitations(physical, mental and temporal) are sort of invisible implicit rules that we all follow in one way or the other. If an entity is not bound by those rules then I don't see why that entity should be treated the same as a human.

Companies already make this differentiation.

For example take captcha and bot detection. Some of the heuristics are based on inherent human limitations like response time, click time, mouse acceleration etc.

I doubt youtube or any other streaming service will be happy if you want to stream all their videos to train a hypothetical human like AI(which views and prepares notes like a human) at a hugely accelerated speed compared to a regular human. You can guess how quickly they will cite fair usage policies.

What I want to say is there are fundamental differences between a human and an AI. So, we should not be quick to dismiss any concerns just because AI can "mimic" humans in certain areas.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#350

Earlier quoted context omitted.

And add to that fact that NYT subscription is hard to unsubscribe from. People have aversion to NYT, even setting aside the bias.

It took me all of 5 minutes to cancel my digital NYT subscription from the following month onward. No idea what you are talking about.

That's only been true for the past few months, and it's been very well documented how complicated the cancelation process used to be [0].

It's funny because I use PayPal for any unknown-to-me site where I don't want to give out my card, but the only site where I've needed their help to cancel something was the New York Times.

[0] https://www.nirandfar.com/cancel-new-york-times/

Post reply on HN