Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

471–480 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#471
post #410

Earlier quoted context omitted.

Creative works like books, TV shows and movies contain facts too.

None of which are copyrightable and infact has been the subject of DMCA abuse like when a Movie uses NASA footage and claims copyright on YouTube videos with the same footage. Copyright is a complex subject, and not as vast as many believe, at the same time ironically it is more vast than i believe it should be. copyright should be much more limiting than it is. Which is at odds with people that believe copyright sho…

Copyright doesn’t exist solely for the “promotion of useful sciences”. https://en.m.wikipedia.org/wiki/Copyright

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#472
post #278

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

Good comment, it was very funny to see how people desperately try to find moral justification for pirating media A but not B. "It's apples to oranges, you see, there are less letters in the NYT article than in the book and they are rendered differently, so it is ok to pirate their work. I did nothing wrong!" :)

I wonder what the reaction of some of the people who browse this forum would be if the output of their careers were so commonly pirated. Somehow, I think most think that this argument doesn't apply.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#473

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

I wouldn't say OpenAI has exactly the same attitude, since they also pulled in thousands of books. Their position has been that it's not piracy, since they don't republish the books; effectively the AI just reads them and learns from them. If GPT can be made to reproduce the original articles, that's a more difficult argument to make.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#474

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

> A sibling comment mentions search engines. I think there's a big difference. A search engine doesn't replace the source, not at all.

Google has been accused for years of replacing sources with their "One Box"--the big answers at the top of the page, which are usually pulled from or corroborated by search results. They don't want you to leave the search results page (where the ads are).

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#475

The NYT is preparing for a tsunami by building a sandcastle. Big picture, this suit won’t matter, for so many reasons. To enumerate a few: 1. Next gen LLMs will be trained exclusively on “synthetic”/public data. GPT-4V can easily whitewash its entire copyrighted training corpus to be unrecognizably distinct (say reworded by 40%, authors/sources stripped, etc). Ergo there will be no copyright material for GPT-5 to reg…

> GPT-4V can easily whitewash its entire copyrighted training corpus to be unrecognizably distinct

Is that just by increasing the temperature, tweaking the prompt, etc.? If you can operate on the raw weights and recreate the original text, copyright infringement still applies.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#476

We developers like to pretend that LLM's are akin to humans and that they've been using things like NYTimes like humans as educational material. But they are not. It's much simpler, proprietary writing is now integrated into the source code of OpenAI, it would be as if I would copy parts of other propriety code and copy paste it into my own codebase. Claiming copy paste is a natural evolving process of millions of ye…

> it would be as if I would copy parts of other propriety code and copy paste it into my own codebase. It's not copy-pasted; it's compressed in a lossy manner. Even GPT4 has nowhere near enough memory to store the entirety of its training data in a non-lossy compression format. Just likes how humans compress the information we read.

If you have a copyrighted photo that I simply put through jpeg compression, am I legally allowed to use that?

Software programs are not humans, and need to be treated differently. Anthropomorphization is one of the slipperiest paths to argue anything.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#477

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

There are two fundamental differences.

First, Open AI is the one doing the pirating here. Hacker News is the host, they aren't doing any pirating or posting any archival links to the copyrighted information themselves.

Second, Open AI charges subscription fees and profits off of the copyrighted material they have pirated, whereas Hackers News does not, nor do the people who post the links.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#479

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

Not quite what parent means, but an interesting angle is: what if you scraped ChatGPT instead.

NYT, or someone's blog? Meh, fair use, and if you say no, you're in the way of progress.

But if you wanted to scrape ChatGPT answers to tweak your network, uh oh, violation of T&C!

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#480

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

Because I'm not interested in the medium itself, as I would be with a Netflix show; I'm not even interested really in the article or the New York Times as an institution. I'm interested in discussing the supposed real-life phenomenon being covered, and the posted content is the primer for that discussion. I think if you get rid of the archive links on HN you need to ban the paywalled content as well. If you want to discuss paywalled content I'm sure you can do that in the article's comment section.
Post reply on HN