Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

611–620 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#611

Earlier quoted context omitted.

Yeah. No one is out there suing the shit out of cliff notes because they published a summary of Catcher in the Rye.

they might if cliff notes starting copy pasting parts of the source into their articles and passing it off as original writing though :)

Newspapers generally don't "pass off" quotes as their own writing. They make clear which parts they quoted.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#612
post #482

Earlier quoted context omitted.

A book, TV show, movie, video game, album, or comic book is not available on the internet served by the copyright holder’s own servers with no authentication or authorization checks. But the NYT is available in that way .

But some are? I believe The Atlantic and The Economist are hard paywalled.

If they're hard paywalled (everyone gets the same login prompt), they won't be available on archive sites.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#613

I think there is a national security aspect to ML models trained on copyrighted data. Countries that allow it will gain a superior technological advantage and outcompete those who disallow training on copyrighted material. I personally believe training LLMs on copyrighted data is copyright infringement if the models are deployed in a way that competes with the copyright holder. But that doesn’t necessarily mean it’s…

You can say the same for any legal enforcement like respecting patent or copyright law or making Champagne outside France. Yet the sky isn’t falling given this reality with so many legally protected industries. Maybe these markets where such an industry might offshore to are too small and insular to be very significant, and are probably language bound to make english models less relevant compared to native language models.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#614

Earlier quoted context omitted.

a court has established this already in japan, where they said anything goes for ai so its best to not to lose a competitive edge with things that people openly publish on the internet, if you put it out there for everyone to see then expect other people to use it

A court in Japan will have no impact on the outcome of a copyright lawsuit in USA. Not to mention that it doesn't really matter how a Japanese court ruled since it's all governed by treaties anyway. They will change their laws if required to.

its not about applying laws across different countries

its about a precedent. If you don't keep up with international competition, you lose.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#615
Isn't the fundamental issue here that the NYT was available in Common Crawl?

If they didn't want to share their content, why did they allow it to be scraped?

If they did want to share their content, why do they care (hint: $88 billion)?

Or is it that they wanted to share their content with Google and other search engines in order to bring in readers but now that an AI was trained on it they are angry?

What wrong thing did OpenAI do specific to using Common Crawl?

Didn't most companies use Common Crawl? Excepting Google, who had already scraped the whole damn Internet anyway and just used their search index?

Is it legal or not to scrape the web?

If I scrape the web, is it legal to train a transformer on it? Why or why not?

To me, this is an incredibly open-and-shut case. You put something on the web, people will read that something. If that is illegal, Google is illegal.

Oh, and do you see the part in the article where they are butthurt that it can reproduce the NYT style?

> "Defendants’ GenAI tools can generate output that recites Times content verbatim, closely summarizes it, and mimics its expressive style, as demonstrated by scores of examples," the suit alleges.

Mimics its expressive style. Oh golly the robots can write like they're smug NYT reporters now--better sue!

It appears that the NYT changed their terms of service in August to disallow their content in Common Crawl[0]. Wasn't GPT-4 trained far before August?

0]: https://www.adweek.com/media/the-new-york-times-updates-term...

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#616

The NYT is preparing for a tsunami by building a sandcastle. Big picture, this suit won’t matter, for so many reasons. To enumerate a few: 1. Next gen LLMs will be trained exclusively on “synthetic”/public data. GPT-4V can easily whitewash its entire copyrighted training corpus to be unrecognizably distinct (say reworded by 40%, authors/sources stripped, etc). Ergo there will be no copyright material for GPT-5 to reg…

> rent seeking media companies Rent seeking? Media companies that actually create content are rent seeking? Versus the garbage hallucinations AI creates?

Rent seeking is an awful term that was from the beginning intended to describe anyone pursing a political or legal goal that deviates from a pure free market economy. As Econlib writes:

> ”Rent seeking” is one of the most important insights in the last fifty years of economics and, unfortunately, one of the most inappropriately labeled. Gordon Tullock originated the idea in 1967, and Anne Krueger introduced the label in 1974. The idea is simple but powerful. People are said to seek rents when they try to obtain benefits for themselves through the political arena. They typically do so by getting a subsidy for a good they produce or for being in a particular class of people, by getting a tariff on a good they produce, or by getting a special regulation that hampers their competitors. Elderly people, for example, often seek higher Social Security payments; steel producers often seek restrictions on imports of steel; and licensed electricians and doctors often lobby to keep regulations in place that restrict competition from unlicensed electricians or doctors.

https://www.econlib.org/library/Enc/RentSeeking.html

This is linked in the wikipedia article, which is even more confused:

https://en.wikipedia.org/wiki/Rent-seeking

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#617

Earlier quoted context omitted.

> You’re not paying to enjoy the content, you’re paying to experience the content. Not sure about others, but I'm not.

Would you make the same argument for a sporting, theatrical or music event? That you should be refunded if you didn't enjoy it?

Does it matter? Sounds to me like an apples and oranges comparison.

If I read an article in the NYT then I'm paying for what I took away from it, not for the amount of time that it allowed me to kill.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#619
post #198

Under existing condition an AI news site seems like a good investment idea. Its AI could read all relevant news sources and retell them and republish them in its own articles. It could even have its own AI editors and contributors. Cannot see how human news companies could compete.

>Cannot see how human news companies could compete. News ultimately comes from physical sources on the ground, which currently AI has no way of doing.

That style of journalism is nearly dead. True on the ground investigative journalism is hardly done today, most is just reporting existing public information releases. You don’t have to be at the presser when everything the police chief says will be put in an online transcript.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#620
post #373

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

> Is that fair use? As always, the answer is.. "it depends". I guess it depends mostly on the jurisdiction that applies to you. "Fair use" can have rather different legal meaning (or not exist at all) in different countries.

Fair use is specific to the US, as far as I'm aware. Moreover, Congress had to codify fair use (turn fair use common into statutory law in the form of 17 U.S. Code § 107) in order to make copyright statutes compatible with the First Amendment. Most other countries don't have freedom of expression and freedom of the press, so copyright law in a different country usually lacks a unifying exception test like fair use to supplement the specific enumerated exceptions.
Post reply on HN