Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

441–450 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#441
post #318

Earlier quoted context omitted.

There's no way to get your money back if you didn't like the content. If they don't want their articles to be read for free then they should keep them out of my view. And certainly not use clickbaity headlines. Information can be copied and they should accept it, or change their business/distribution model.

So if I went to a cinema and didn't like the movie, I should be entitled for a return, right? Or if I went into a museum and didn't like the art displayed there? If you are advocating for a free for all libertarian dystopia, well, I have some bad news for you - they never work.

I don't agree with the OP but how are refunds a free for all libertarian dystopia?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#442
The thing that bothers me about the whole situation is that OpenAI prohibits using its model output as training data for your own models.

It seems more than a bit hypocritical, no? When it comes to their own training data, they claim to have the right to use any/all of humanity’s intellectual output. But for your own training data, you can use everything except for their product, conveniently for them.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#443

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

The NYT and other newspapers don’t go after the archived link providers. Probably because the newspapers scholarly mission includes things like preservation. But they also have a profit motive or they can’t stay in business.

This implicit permission for the archive links to exist, gives some of us the implicit permission to pirate the content.

Disclaimer: I am a happy subscriber to the NYT (and other digital newspapers).

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#444
post #422

At some point the burden of carrying 100 year old copywriter/patent law will become so onerous a burden on the pace of progress that its enforcement will be antihuman.

It already is, but I don't think this is a good example. NYT has a legitimate case here. They own the material they publish, and GPT-4 is shown to be able to recall entire articles verbatim. That's a violation, clear as day.

The thing about lawsuits is that you make dozens of claims, and the court can rule in favor of some of them, and against others. The question of "is LLM training fair use?" hasn't made it to a high court yet. The court could very easily rule against everything else in the suit.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#445
post #419

Earlier quoted context omitted.

I'm sorry but this is such a bad take. Nice appeal to consequences. In my view, the New York Times is entirely justified in pursuing legal action. They invested time and effort in creating content, only to have it used without permission for monetary gain. A clear violation. Analyzing the factors involved for a "fair use" consideration: Purpose and Character of the Use: While the argument for transformation might hol…

Imo gpt itself is the transformative work.

Ok but it's not

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#446
post #419

The NYT is preparing for a tsunami by building a sandcastle. Big picture, this suit won’t matter, for so many reasons. To enumerate a few: 1. Next gen LLMs will be trained exclusively on “synthetic”/public data. GPT-4V can easily whitewash its entire copyrighted training corpus to be unrecognizably distinct (say reworded by 40%, authors/sources stripped, etc). Ergo there will be no copyright material for GPT-5 to reg…

I'm sorry but this is such a bad take. Nice appeal to consequences. In my view, the New York Times is entirely justified in pursuing legal action. They invested time and effort in creating content, only to have it used without permission for monetary gain. A clear violation. Analyzing the factors involved for a "fair use" consideration: Purpose and Character of the Use: While the argument for transformation might hol…

I don’t think the original point being made was that NYT wasn’t justified in bringing the action. The point that was being made was the suit would be ultimately meaningless in the long term even if it was successful in the short term. There is a potentially more significant risk in the future that this suit will not protect against because of the reasons enumerated by the author. While the author is speculating, the law struggles with technology and adapting to change, which makes their prediction useful because it does highlight the problems that are coming that can’t be readily mitigated through legal precedent.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#447
post #390

Earlier quoted context omitted.

You’re not paying to enjoy the content, you’re paying to experience the content. And as long as you had the opportunity to experience the content, you’ve gotten what you paid for. I don’t see “I don’t like it” as a valid reason for a refund.

> You’re not paying to enjoy the content, you’re paying to experience the content. Not sure about others, but I'm not.

Your personal opinion on the matter has little weight here.

It doesn't matter what you think you're paying for or should be paying for, the fact of the matter is that you're paying for the effort people put in bringing that to you. So you are, whether you want to be or not.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#448

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

The difference is that an individual pirating news is simply reading the article. OpenAI intends to digest news articles to the point of packaging them and reselling.

My uncle used to distribute daily newspapers and his saying was "News ages like a fish".

OpenAI is allegedly using NYTimes articles to train a computer and sell its services. I see different use scenarios.

I guess another way to look at it is that human just reads the pirated material. A computer makes a verbatim copy and analyzes it to the point to mimicry and sells fuzzy versions.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#449
post #435

Earlier quoted context omitted.

Scraping is legal, and this seems like a transformative work to me.

Returning the full text of an article verbatim seems to me like the opposite of "transformative."

In the screenshot for the article you can see that the LLM says it is "Searching for: carl zimmer article on the oldest DNA". That, and what I know about how LLMs work, suggest to me that rather than the article being stored inside the trained LLM it was instead downloaded in response to the question. So the fact that the system is providing the full text of the article doesn't really go to whether training the LLM is a transformative use or not.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#450

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

I believe it's tolerated here based on the site guidelines. I have always thought this was the case because otherwise these posts would all be pay to play which would limit who could participate and turn HN into more of a subscription farm. Maybe the way to make everyone feel ok about it is to disallow links to paywalled content.
Post reply on HN