Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

361–370 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#361
post #194
post #72

Would be funny if NT Times won this and all commercial LLMs were shut down. Then LLMs would be distributed only via torrents, like most copyright infringing media.

What will happen in this case is that large content providers will get paid directly and smaller content providers will get rolled up into a licensing bag and get small indirect payouts. For example, we might see a model where people who's books have been used will get a pay out proportionate to the sales of the book (perhaps), so if your books sells just a few thousand copies expect $20 but if you sell millions expe…

> large content providers will get paid directly

I'm sure that's what they want, but I'm not sure that's what the outcome will be. What if they want to charge a prohibitive amount of money for their content?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#362

Earlier quoted context omitted.

I would be "happier" to pay a subscription to an aggregation platforms like hackernews or reddit to access archived articles that are linked to these sites. In turn a proportion of that could be passed on to the underlying publishers that I actually visit. I have nearly zero interest in reading articles that aren't linked to from an aggregation site. I don't want to read theguardian.com, or nytimes.com, or washington…

This is a common statement, but every attempt to sell that service has been a dismal failure. See for example blendle.

Blendle failed because they went into competition with the papers whose content they reproduced.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#363

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

> I think that's something worth reflecting on, about why we feel it's OK to pirate news articles, but not other IP

As you noted it is not the norm to post pirate links here for IP other than news articles, but that doesn't mean that a lot of people think it is not OK to pirate those other forms of IP.

In nearly any big discussion that even remotely involves video streaming there will be numerous posts from people explaining why they pirate (usually with ridiculous justifications like "subscribing is not an option because even though this paid service does exactly what I want now at a price that is trivial for me they might someday later change").

The impression I've gotten is that piracy of nearly everything is widely felt to be OK here. Information wants to be free, yada yada.

About the only piracy that is consistently frowned upon here is piracy of open source software. When some company sells an embedded device that uses GPL code without releasing the corresponding source that's viewed as just a little short of a crime against humanity.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#364
post #265

Earlier quoted context omitted.

To me, there is a sense that the news, which is real information about the society that we currently live in, should be availabe to all participants of that society. The notion of being a good citizen requires that one stays informed. Books, movies, videogames etc. don't have that role and are more consumption goods.

> should be available to all participants of that society. Who pays?

There's a few possible models here:

Public donors ALA Patreon

People doing it in their free time because they care a lot about the subject (nowadays with things like Twitter its quite possible for an independent obsessive to write a good piece on, for instance, the Ukraine War by mostly referring to open sources and public announcements by governments and corporations)

Government sponsorship ala BBC

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#365

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

Funny, I don't see it as a moral thing but more a "what can you get away with" thing. I fully assume that if I was to post a magnet link to a torrent for whatever the link was about, I would be banned. Morally speaking, I think it's perfectly reasonable to download a copy of something and either read the relevant info for my current task or to sample it to decide if I want to buy it. I see it no different to using th…

So downloading a movie from piratebay is no different to using the library?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#366

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

> , a movie, a video game, an album, a comic book, or any other form of IP, it is in fact very much _not_ the norm for the top-rated comment to be a Pirate Bay link. Probably because most print media is garbage and nobody in their right mind would actually pay to read them

> Probably because most print media is garbage and nobody in their right mind would actually pay to read them

NYTs revenue keeps growing though.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#367

Earlier quoted context omitted.

Yeah. No one is out there suing the shit out of cliff notes because they published a summary of Catcher in the Rye.

they might if cliff notes starting copy pasting parts of the source into their articles and passing it off as original writing though :)

The Tolkien estate should get busy suing all the fantasy writers, comic artists, game developers and board and card game companies. Lots of cash there.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#368
I think there is a national security aspect to ML models trained on copyrighted data. Countries that allow it will gain a superior technological advantage and outcompete those who disallow training on copyrighted material. I personally believe training LLMs on copyrighted data is copyright infringement if the models are deployed in a way that competes with the copyright holder. But that doesn’t necessarily mean it’s something we should disallow.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#369

Earlier quoted context omitted.

It is actually pirating content by companies for humongous profit, or pirating by individual human beings for free access to culture and entertainment, oftentimes for content one has already paid for, but rendered inaccessible by megacorporations.

Which content making businesses earn humorous profit margins? Are all the journalist layoffs a fever dream? This is one of the more profitable ones, and only because they employ unscrupulous tactics: https://www.macrotrends.net/stocks/charts/NWS/news/profit-ma... This is NYT, the most successful news business: https://www.macrotrends.net/stocks/charts/NYT/new-york-times... As for movies/tv show/music makers, let’s ju…

> Which content making businesses earn humorous profit margins?

You got my point backwards: AI companies will make it from the pirated content, that individual users don't make.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#370

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

To me, there is a sense that the news, which is real information about the society that we currently live in, should be availabe to all participants of that society. The notion of being a good citizen requires that one stays informed. Books, movies, videogames etc. don't have that role and are more consumption goods.

> which is real information about the society that we currently live in, should be availabe to all participants of that society.

Who should pay the journalists or the investigative reporters?

Post reply on HN