It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
NY Times copyright suit wants OpenAI to delete all GPT instances
341–350 of 921 posts
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#342The suit demonstrates instances where ChatGTP / Bing Copilot copy from the NYT verbatim. I think it is hard to argue that such copying constitutes "fair use". However, OAI/MS should be able to fix this within the current paradigm: Just learn to recognize and punish plagiarism via RLHF. However, the suit goes far beyond claiming that such copying violates their copyright: "Unauthorized copying of Times Works without p…
> Just learn to recognize and punish plagiarism via RLHF. I'm not sure how your proposal would actually work. To recognize plagiarism during inference it needs to memorize harder. Kinda funny if it works though. We'd first train them to copy their training data verbatim , then train them not to. That is how it works, right? They're trained to copy their training data verbatim because that's the loss function. It's ju…
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#343Earlier quoted context omitted.
Possibly because once an article is published the author receives no further payment. In all other mediums, there are residuals and royalties to be paid to the creators of the work.
Articles have ads on them, how are they not residual payments based on views?
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#344It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#345Earlier quoted context omitted.
I would be "happier" to pay a subscription to an aggregation platforms like hackernews or reddit to access archived articles that are linked to these sites. In turn a proportion of that could be passed on to the underlying publishers that I actually visit. I have nearly zero interest in reading articles that aren't linked to from an aggregation site. I don't want to read theguardian.com, or nytimes.com, or washington…
This is a common statement, but every attempt to sell that service has been a dismal failure. See for example blendle.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#346Earlier quoted context omitted.
It is actually pirating content by companies for humongous profit, or pirating by individual human beings for free access to culture and entertainment, oftentimes for content one has already paid for, but rendered inaccessible by megacorporations.
Which content making businesses earn humorous profit margins? Are all the journalist layoffs a fever dream? This is one of the more profitable ones, and only because they employ unscrupulous tactics: https://www.macrotrends.net/stocks/charts/NWS/news/profit-ma... This is NYT, the most successful news business: https://www.macrotrends.net/stocks/charts/NYT/new-york-times... As for movies/tv show/music makers, let’s ju…
If only piracy would actually harm these businesses but alas as often demonstrated it has zero effect on their bottom line, if anything it increases their profits.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#347It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
Good comment, it was very funny to see how people desperately try to find moral justification for pirating media A but not B. "It's apples to oranges, you see, there are less letters in the NYT article than in the book and they are rendered differently, so it is ok to pirate their work. I did nothing wrong!" :)
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#348Let's say OpenAI was trained on all the Windows source code (without approval from MS).
GPT could pretty much replicate the windows code with even not that clever prompt by any user. "Write an OS CreateProcess function like Windows 10 source code would have."
It would infuriate MS to put it mildly, enough to start a lawsuit.
I know the license to the MS source code and NYT articles aren't the same.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#349Earlier quoted context omitted.
Another factor to consider is that neural nets can function as lossy compression, which becomes extremely evident when using models that are overfit. Sometimes they're so overfit that the compression isn't even lossy, and the data is encoded verbatim in the NN.
Yes, but this then hits against learning/understanding and compression being fundamentally the same thing . I can't think of a better way to argue in favor of "it's fine if human does it, therefore it's fine if LLM does it", than from the "lossy compression" angle.
Our collective human limitations(physical, mental and temporal) are sort of invisible implicit rules that we all follow in one way or the other. If an entity is not bound by those rules then I don't see why that entity should be treated the same as a human.
Companies already make this differentiation.
For example take captcha and bot detection. Some of the heuristics are based on inherent human limitations like response time, click time, mouse acceleration etc.
I doubt youtube or any other streaming service will be happy if you want to stream all their videos to train a hypothetical human like AI(which views and prepares notes like a human) at a hugely accelerated speed compared to a regular human. You can guess how quickly they will cite fair usage policies.
What I want to say is there are fundamental differences between a human and an AI. So, we should not be quick to dismiss any concerns just because AI can "mimic" humans in certain areas.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#350Earlier quoted context omitted.
And add to that fact that NYT subscription is hard to unsubscribe from. People have aversion to NYT, even setting aside the bias.
It took me all of 5 minutes to cancel my digital NYT subscription from the following month onward. No idea what you are talking about.
It's funny because I use PayPal for any unknown-to-me site where I don't want to give out my card, but the only site where I've needed their help to cancel something was the New York Times.