Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

281–290 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#281

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

The archive link doesn't threaten their jobs and helps them avoid paying for NYT. It's NIMBY, or rather it's true form of NIIIM (Not if it impacts me).

Hypocrites are EVERYWHERE and are the majority.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#282

Earlier quoted context omitted.

Another factor to consider is that neural nets can function as lossy compression, which becomes extremely evident when using models that are overfit. Sometimes they're so overfit that the compression isn't even lossy, and the data is encoded verbatim in the NN.

Yes, but this then hits against learning/understanding and compression being fundamentally the same thing . I can't think of a better way to argue in favor of "it's fine if human does it, therefore it's fine if LLM does it", than from the "lossy compression" angle.

> fundamentally the same thing

I fundamentally disagree. That's not some established fact, just a narrative used by those who wish to plagiarize using "AI".

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#283

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

It is am ethical grey area, but if the paywall applied to all user agents, which would make it similar to say buying a Kindle book, then you might see that as pirating, whereas if you use an archive service that was served the HTTP response and cached it, then you are using a proxy UA. If the news/magazine doesn't want this they can simple serve a cut down or zero length article to all non-paying viewers! But they wa…

We can extend this analogy. What if someone put up a proxy, that has a legal Netflix subscription and which "watches" streams of Netflix shows, captures actual RGB values of pixels and re-streams the resulting video to anyone else? Isn't it the same "proxy" excuse?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#284
post #278

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

Good comment, it was very funny to see how people desperately try to find moral justification for pirating media A but not B. "It's apples to oranges, you see, there are less letters in the NYT article than in the book and they are rendered differently, so it is ok to pirate their work. I did nothing wrong!" :)

It is actually pirating content by companies for humongous profit, or pirating by individual human beings for free access to culture and entertainment, oftentimes for content one has already paid for, but rendered inaccessible by megacorporations.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#285

Earlier quoted context omitted.

Possibly because once an article is published the author receives no further payment. In all other mediums, there are residuals and royalties to be paid to the creators of the work.

And add to that fact that NYT subscription is hard to unsubscribe from. People have aversion to NYT, even setting aside the bias.

It took me all of 5 minutes to cancel my digital NYT subscription from the following month onward. No idea what you are talking about.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#286

Earlier quoted context omitted.

To me, there is a sense that the news, which is real information about the society that we currently live in, should be availabe to all participants of that society. The notion of being a good citizen requires that one stays informed. Books, movies, videogames etc. don't have that role and are more consumption goods.

I agree, but nothing worth having is free. NYT and other news outlets have to ultimately pay reporters to go out into the world and do the work. The reporters are not priests, and the NYT is not a church that lives off donations and tax exemptions. They need money to operate, and you may disagree with how they try to collect that money (paywall) but that doesn't solve their funding problem. How would you pay for news…

> How would you pay for news otherwise?

You could subsidise news via "public service" style stipends. Much like having a government owned "independent" news service (eg the BBC) this comes with a high risk of corruption. Don't bite the hand that feeds and all that.

You could implement a much lower friction non-recurring payment system. I'd be far more tempted to drop a little money on a fixed term (5 articles, 1 day, ???) setup than a subscription.

Realistically, I am not paying for more than 1 long running sub. And there are > that number of solid outlets.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#287

Companies that have content all see dollar signs. NYT won't mind if you use their content to train LLMs - as long as they get a commission. Reddit will shut down their free API and make you pay to get training content. Discord is going to be selling content for AI training too - if they haven't already done so. Twitter is doing it. They didn't care before because LLMs were just experiments. Now we're talking trillion…

"They" also include the people working there. Why someone work with full time writing articles should give the work for free just let someone to train it and make money out of it as a consequence?

You understand that news aren't copyrightable right?

You're fighting a scarecrow that doesn't exist...

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#288

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

Funny, I don't see it as a moral thing but more a "what can you get away with" thing.

I fully assume that if I was to post a magnet link to a torrent for whatever the link was about, I would be banned.

Morally speaking, I think it's perfectly reasonable to download a copy of something and either read the relevant info for my current task or to sample it to decide if I want to buy it. I see it no different to using the library or browsing at a book store.

Perhaps once news organisations can work out how to effectively wield the DMCA hammer against archive links we'll see the practice of posting them stop.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#289
post #265

Earlier quoted context omitted.

To me, there is a sense that the news, which is real information about the society that we currently live in, should be availabe to all participants of that society. The notion of being a good citizen requires that one stays informed. Books, movies, videogames etc. don't have that role and are more consumption goods.

> should be available to all participants of that society. Who pays?

The government (thus the people, in a so called sharing of public burden)!

For example in Hungary there is an official news agency ran by the government, with (cumbersome) free access for everybody. Of course this does provide somewhat biased presentation of some facts, but on many topics it provides unbiased access to news for any citizen.

This is actually pretty common in Europe, often funded by mandatory fees (for some reason not branded as taxes) certain appliance owners need to pay (UK TV license, German Rundfunkbeitrag). For this fee people get access to news and cultural programmes for free via different media (radio, TV, internet).

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#290

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

This is only an interesting juxtaposition if you have fully internalized and accepted the myth of people and corporations being interchangeable.
Post reply on HN