Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

571–580 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#571

Earlier quoted context omitted.

This seems very false to me. Spotify is the prime example. They offer a good product that covers a 100% of my needs at a reasonable price. If that was an option for say UFC or engineering books, you bet I’d be subscribed. But being forced to read through some crappy reader software when I need the book source to take annotations in another software doesn’t work, so here we are. Same with the absurd pay per view busin…

For books, if it's a client reader software frustration, then you should still buy the digital version and then you can pirate the PDF book and use as desired within the constraints of copyright law (e.g. don't go sharing the PDF). That way you get the client you want but you still paid the content creator. But to use the argument, "oh, I don't like their client so I'm going to not pay them" is BS. For UFC, your comp…

The problem with buying by the crappy DRM version is that it provides no incentive to the publisher to change. I have thought about this long and hard, but ultimately the only way Spotify came about was because nobody bought the terrible DRM’d music the labels wanted to foist on us. We need to inflict the same pain for books. Personally, I think it would be preferable to donate the same amount to the Books Trust or your local library.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#572
post #390

Earlier quoted context omitted.

You’re not paying to enjoy the content, you’re paying to experience the content. And as long as you had the opportunity to experience the content, you’ve gotten what you paid for. I don’t see “I don’t like it” as a valid reason for a refund.

> You’re not paying to enjoy the content, you’re paying to experience the content. Not sure about others, but I'm not.

Would you make the same argument for a sporting, theatrical or music event? That you should be refunded if you didn't enjoy it?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#573

NYT's perspective is going to look so stupid in future when we put LLMs into mechanical bodies with the ability to interact with the physical world, and to learn/update their weights live. It would make it completely illegal for such a robot to read/watch/listen to any copyrighted material; no watching TV, no reading library books, no browsing the internet, because in doing so it could memorise some copyrighted conte…

Are those LLMs independant citizens we are going to give rights to? Then I'm fine with that. Are they all owned by one mega-corporation, which is going to do as capitalism does, and use them to squeeze money out of all of us? Then I'm happy to ban them.

"Let's ban something capable of diagnosing medical conditions and letting coma patients to communicate with an EEG because it learned the relationships between words from a giant data set of scraped data and is owned by a company" is a pretty callous take IMO.

The opportunity cost of holding this technology back is going to literally be millions of people's lives given current trends in its emerging applications.

Police usage, not training.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#574
post #352

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

> Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. [...] And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or any other form of IP, it is in fact very much _not_ the norm for the top-rated comment to be a Pirate Bay link. If the sto…

So, what allows accessing content under IP illegally is not liking the marketing strategy of the content owner?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#575

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

If ChatGPT is based on neural networks, with no actual save-and-replicate facsimile behaviour, it no more "copies" original work than I do when I tell you about the news article I read today. I'd say the only real reason the Piratebay links thing you mentioned is not the norm is purely because those media sources have done a better job of striking fear into people doing that, so it's gone more underground. I.e. they'…

> it no more "copies" original work than I do when I tell you about the news article I read today

When you tell people about some news article you read earlier you repeat it exactly verbatim? You also give this out to potentially millions or hundreds of millions of people for commercial purposes?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#576

Earlier quoted context omitted.

People used to leave newspapers in the trash, on the train, all over the place. Anyone could pick them up and read for free. I think it's reasonable for folks to carry this attitude into the digital age. People feel like news is something to share, it's not the source of creative expression, it's facts and as such we feel entitled to know the facts about our world and what is happening that might affect us.

That newspaper was likely paid for by someone, and could only be read by one person at a time.

And what if the person picking up the paper would stand up and shout the content of the article so all the people on the train would hear?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#577

NYT's perspective is going to look so stupid in future when we put LLMs into mechanical bodies with the ability to interact with the physical world, and to learn/update their weights live. It would make it completely illegal for such a robot to read/watch/listen to any copyrighted material; no watching TV, no reading library books, no browsing the internet, because in doing so it could memorise some copyrighted conte…

I disagree. The verbatim part is the problem. You’re drawing a comparison to how humans operate except we’re not allowed to operate like that. While harder to do as a human, if memorised a copyrighted book and then did a live reading on TV, or produced replicas from memory and sold them (the most comparable example), I’d be sued. Humans produce derivative work all the time, and it’s fine for LLM’s to do that, but you…

Then we should be focused on policing the usage of the model, not the training of it.

That's the point at which infringement occurs in your example. It's not the memorizing that's the infringement, it's the reproduction from your memory.

We shouldn't be regulating your hippocampus encoding the book, but your reproducing the book from that encoding.

Similarly, we shouldn't be regulating the encoding of material into the NN, but the NN spitting back out the material.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#578
Summarizing the article: The most damning thing here is the "ChatGPT as a search engine" feature, which appears to run an agent which performs a search, visits pages, and returns the best results.

In doing this, it is bypassing the NY Times paywall, and you can read full articles from today by repeatedly asking for the next paragraph.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#579

Earlier quoted context omitted.

Historically newspapers leaned more on competition law than copyright, because their pages are supposed to be filled with non-copyrightable facts.[1] Copying part, but not all, of a factual article, significantly after the relevant event, was considered to be a promotion (not unfair competition) and a nice thing to do for the journalists. Things change, people lose sight of the original principles. [1] https://en.m.w…

> their pages are supposed to be filled with non-copyrightable facts This is rather inaccurate. A fact is Hitler invades Poland. You're right, nobody can copyright this idea, as it is just a fact. However, if I then write a 500-word article describing the scene of Hitler invading Poland, have short quotes from some civilians there, etc. that particular arrangement of ideas and words is copyright. AP can't go and sue…

Yes, the prose was always under copyright, but the key point for the case linked in the wikipedia article is:

> INS members would rewrite the news and publish it as their own without attribution to AP.

So the case hinged on INS indeed reporting facts that differed in exposition.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#580

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

What you described is entirely fair use, actually.

Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost all of them are directly sourced from NYT and others. They'll say "so and so happened, according to The Times" (and usually link the article there)

Post reply on HN