Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

71–80 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#71

Earlier quoted context omitted.

No it's the death of the corporate content hosting web - the open web was never about making money with your blog post/irc chat/usenet group/etc content, at least in my opinion. Let data be free! If someone wants to use it to make money, well, it's open, just like open source. It's still not okay to take open source work and claim it as your own, which is what copyright should be limited to. Stealing a photo or plagi…

Maybe I don’t want my non-corporate art to be a part of some large corporation’s training data.

User-agent: GPTBot Disallow: /

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#72
post #9
post #3

What’s going to be the name used for the laws that attempt to tackle machine paraphrasing?

Paraphrasing is not the issue. The issue is that OpenAI copied the Times ’ creative works into a GPU to train a model. That copy was likely neither licensed nor fair use.

Google does the same to produce a search index.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#73
post #71

Earlier quoted context omitted.

Maybe I don’t want my non-corporate art to be a part of some large corporation’s training data.

User-agent: GPTBot Disallow: /

That was my gut reaction too, but presumably unless it becomes regulated, at least some competitors to OpenAI won't respect any robots.txt and thus any open content might be training data.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#74
post #71

Earlier quoted context omitted.

Maybe I don’t want my non-corporate art to be a part of some large corporation’s training data.

User-agent: GPTBot Disallow: /

ATMO you shouldn't have to maintain knowledge of what kind of crawler bot exist and having to maintain deny list. It should be the opposite, only expressedly allowed content should be crawled by mainaining allow lists.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#75
post #49

I don't think it's an exaggeration to say that LLMs might lead to the end of the open web, or at least a drastically reduced version of it. So much of these model's utility is in directly competing with the producers of the training data. Content creators and aggregators are seeing more and more reason to restrict and limit access, to avoid having AI companies consume all of their data and then be the ones making mon…

Let's assume that happens. How do I hedge against it? Is there a convenient way to mirror the bits of the web that are open now? Perhaps a mirror of archive/WayBack machine that could be viewed locally similar to Wikipedia dumps?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#76
post #49

I don't think it's an exaggeration to say that LLMs might lead to the end of the open web, or at least a drastically reduced version of it. So much of these model's utility is in directly competing with the producers of the training data. Content creators and aggregators are seeing more and more reason to restrict and limit access, to avoid having AI companies consume all of their data and then be the ones making mon…

Really? Seems like it's not so different conceptually to search engines. They also make money by indexing "training data" of a sort. But the open web trundles on regardless, because the search engines found ways to cut the content producers in on it. I see no reason why that can't be the case here too, with AI companies training their models to act more like search engines when data comes from certain sources - i.e. they'll point you in the right direction but not directly answer you. That will suck for the AI users, but, for many questions the right answer will be available in open or bulk licensable training sets anyway. For example you can get legal access to nearly all books by doing deals with publishers, as Google Books has demonstrated, you can get access to map data by generating it yourself or doing deals with digital map companies etc.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#77
post #70

There is a very real risk that we end up with an inferior product cannibalizing a superior one and driving it out of business. Moreover, AI would seem to be even more susceptible to capture and manipulation than conventional media. When it's a question of guiding thought I prefer the humanities to tech. (Same with art.)

> we end up with an inferior product cannibalizing a superior one and driving it out of business.

In case that print is meant by inferior product: The same argument could've been brought up for Napster, where traditional distribution via CD printing through music labels are the inferior product driving the superior one out of business. Or rather it's big labels suing Napster out of business.

I also hold a dislike for the copyright lobby, but this matter is serious. The question of whether the training of LLM's with copyrighted data is a legitimate one, as OpenAI did not just use contributions from large media outlets like NYT but capitalized on small contributions from individual contributors.

A ruling in favor of copyright would force OpenAI to shut down - but given their impressive demo of the tech, I hope we'd see more open and accessible versions of these models emerge.

As impressive as ChatGPT is, I dislike having my access to information governed by some large corporate entity. I also dislike a company directly capitalizing on my contributions without my explicit consent.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#78

Earlier quoted context omitted.

Did the Times grant a license to every router on the internet to transmit its intellectual property to other routers? If not, the judge should grant an injunction contingent on requiring the Times to verify that every person who accesses their content is doing so only over routers and other devices with express written authorization, for every step in the process. Maybe even extend it to browsers and client libraries…

You don’t need a fair use exemption for transient copies in service of licensed or fair uses. Computers and networks have been around a long time. These issues have been given a good workout.

But the copy used during training is itself transient, so all this boils down really to the question of whether training a machine is fair use. Which can't be answered here exactly because the concept of fair use is deliberately vague, so this will boil down to a lawsuit and probably go to the Supremes. The USA will work something out that's reasonable as they always do and, lacking AI companies and often the concept of fair use to begin with, the rest of the world will never work out the necessary case law and fall even further behind. Possible exception: UK, which does have some significant LLM companies and also fair use law, albeit as is often the case with the UK these firms are US/UK hybrids with significant presence in both countries and the legal HQ is usually in the USA.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#79

>If, when someone searches online, they are served a paragraph-long answer from an AI tool that refashions reporting from The Times, the need to visit the publisher's website is greatly diminished, said one person involved in the talks. If, when someone reads a newspaper, they are served a paragraph-long answer from an NYTimes reporter that refashions reporting from local sources, the need to interact with the local…

Not to mention there are already entire news outlets that exist to summarize paywalled journalism. Ever read an article that starts with "The NYTimes reports..."?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#80
post #70

There is a very real risk that we end up with an inferior product cannibalizing a superior one and driving it out of business. Moreover, AI would seem to be even more susceptible to capture and manipulation than conventional media. When it's a question of guiding thought I prefer the humanities to tech. (Same with art.)

> we end up with an inferior product cannibalizing a superior one and driving it out of business. In case that print is meant by inferior product: The same argument could've been brought up for Napster, where traditional distribution via CD printing through music labels are the inferior product driving the superior one out of business. Or rather it's big labels suing Napster out of business. I also hold a dislike for…

[dead]
Post reply on HN