Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

641–650 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#641
post #493
post #195

Earlier quoted context omitted.

sure, but if I use an LLM to write a novel/article, I can be sued in civil court not the LLM. but, more importantly, OpenAI can also be sued for tortious interference? (basically the civil equivalent of accessory)

> sure, but if I use an LLM to write a novel/article, I can be sued in civil court not the LLM That's function of the legal system, not of the technology. If tomorrow someone made a perfect dolphin-Esperanto translator and proved Dolphins were as smart as humans, you still can't sue a dolphin until the legal system says so.

Wouldn't you find out by suing the dolphin and seeing if it holds up in court?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#642
post #419

Earlier quoted context omitted.

I'm sorry but this is such a bad take. Nice appeal to consequences. In my view, the New York Times is entirely justified in pursuing legal action. They invested time and effort in creating content, only to have it used without permission for monetary gain. A clear violation. Analyzing the factors involved for a "fair use" consideration: Purpose and Character of the Use: While the argument for transformation might hol…

> it's clearly not helping their market value if people are checking on ChatGPT instead of reading a NYT article. People are not using ChatGPT as a replacement for current news, and because of hallucinations, no one should be using it for past news either. I wouldn't remotely call ChatGPT a competitor of NYT traffic, like I would Reuters or other news outlets.

The intended result is clearly to supplant other information sources in favor of people getting their information from ChatGPT. Why should it matter to legality that the tech isn't good enough for the goal?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#643
post #625

Earlier quoted context omitted.

> To keep things simple, let's say I never regurgitate chunks of verbatim NYT articles, maybe quite short snippets. You just described Google. When you think about it, it's surprising that Google is legal. However, it is well established that what Google does is perfectly legal. Remember that internally Google keeps and uses complete verbatim copies of every web page they index. Yes, Google offers a link to the sourc…

Any publisher can opt out of google. Publisher also have substantial control over titles and snippets shown in google, whether an article appears in google news, etc Paraphrasing is also known as cloning and is often a copyright violation

Copyright law doesn't mention opt outs or search engine snippet controls. It's not clear to me that robots.txt is the singular thing that makes Google legal.

In US copyright law facts cannot be copyrighted, so copyright on factual content like newspaper articles is limited. Simply replacing a few words wouldn't work, but I am certain that GPT-4 is capable of paraphrasing factual content at a level that would not be considered infringement if a human did it.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#644

Earlier quoted context omitted.

Yes, but this then hits against learning/understanding and compression being fundamentally the same thing . I can't think of a better way to argue in favor of "it's fine if human does it, therefore it's fine if LLM does it", than from the "lossy compression" angle.

Is there some LLM meta where understanding and compression are argued to be the same thing I’m not aware of? Anyone got more details on this? Superficially it sounds like total BS; a highly compressed zip file does not exhibit any characteristics of learning. Algorithmically derived highly compressed video streams do not exhibit characteristics of learning. ? I’ve vaguely heard the learning can be considered to exhib…

The idea precedes LLMs by a couple of decades and is thought to apply more broadly within ML/AI than being a specific meta for LLMs. http://prize.hutter1.net/ has been around for a while, there is a link in there to the earlier work (called AIXI?).

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#646

Earlier quoted context omitted.

> Gradient descent and backpropagation don't take place in the brain. Not exactly, no, but the 'neurons that fire together wire together' way of learning has a pretty similar effect. > LLMs "learn" in the same way that Excel sheets "learn". I've never seen an excel sheet do anything like backpropagation.

> I've never seen an excel sheet do anything like backpropagation. Not strictly in the sense you mentioned (assuming that you mean "by themselves") but people may find [1] and [2] interesting. [1] https://pub.towardsai.net/building-a-neural-network-with-bac... [2] https://towardsdatascience.com/demystifying-feed-forward-and...

Sadly, I have seen one. It was a vba script from the late 90s that used a simple dense multilayer network to do some unsupervised pattern classification. The linear algebra tools in vba/excel along with the solvers are all native dll code and the vba itself is all AOT compiled to native, so it typically runs very fast, and for small matrices it beats out numpy by an order of magnitude due to the ffi overhead. Was it the wrong tool? It depends on your constraints, but probably. It did work though.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#647

Earlier quoted context omitted.

> it would be as if I would copy parts of other propriety code and copy paste it into my own codebase. It's not copy-pasted; it's compressed in a lossy manner. Even GPT4 has nowhere near enough memory to store the entirety of its training data in a non-lossy compression format. Just likes how humans compress the information we read.

If it’s lossy compressed how come they have verbatim content from NYT in there that’s easy to recall? That’s what the lawsuit is about.

Many humans have photographic memories. Not common, but not unheard of for people to be able to memorize long portions of text verbatim.

For example, the Wikipedia article

https://en.wikipedia.org/wiki/List_of_people_claimed_to_poss...

contains several examples of people who were able to look at pages and recite them back. That is actually a much stronger ability than GPT since GPT has presumably looked at them 100 times.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#648

Earlier quoted context omitted.

The movie/tv show and music business can keel over and die tomorrow - it wouldn’t affect the value of art produced by humans at all. I see those more as exploitative leeches than as contributing anything positive. If only piracy would actually harm these businesses but alas as often demonstrated it has zero effect on their bottom line, if anything it increases their profits.

What do you mean by "art"?

Hard question, but in the context of my comment I would say any kind of visual media or music

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#649
post #630
post #615

Isn't the fundamental issue here that the NYT was available in Common Crawl? If they didn't want to share their content, why did they allow it to be scraped? If they did want to share their content, why do they care (hint: $88 billion)? Or is it that they wanted to share their content with Google and other search engines in order to bring in readers but now that an AI was trained on it they are angry? What wrong thin…

If you read the complaint, it explains this pretty well. The use of copyrighted content by search engines is fundamentally different from the way LLMs use that same content. The former directs traffic (and therefore $$) to the publisher, the latter keeps the traffic for itself. The legal misconception I want to flag in your logic is the notion that all uses of the Common Crawl are equally infringing/non-infringing. I…

From my original comment:

> Is it legal or not to scrape the web?

> If I scrape the web, is it legal to train a transformer on it? Why or why not?

At no point did I say anything about hosting a mirror of the NYT website, with free articles. Obviously. Because OpenAI didn't do that. Some NYT lawyer tried to get ChatGPT to write a NYT article. Maybe first they should have actually done a Google search and shut down some of the actual content farms which simply copy NYT content such as [0]. But instead, we get this.

[0]: https://salaminv.com/news_file/

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#650

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

But is it legal for me to read the NY Times about a war, and then charge people to interview me as an "expert"?
Post reply on HN