Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

481–490 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#481
post #375

Earlier quoted context omitted.

There's a few levels to this... Would it be more rigorous for AI to cite its sources? Sure, but the same could be said for humans too. Wikipedia editors, scholars, and scientists all still struggle with proper citations. NYT itself has been caught plagiarizing[1]. But that doesn't really solve the underlying issue here: That our copyright laws and monetization models predate the Internet and the ease of sharing/paywa…

Can you imagine spending decades of your life, studying skin cancer, only to have some $20/month ChatGPT index your latest findings and spit out generically to some subpar researcher: "Here's how I would cure melanoma!" followed by your detailed findings. Zero mention of you. F-that. Attribution, as best they can, is the least OpenAI can do as a service to humanity. It's a nod to all content creators that they have b…

Can you imagine spending decades of your life studying antibiotics, only to have an AI graph neural network beat you to the punch by conceiving an entire new class of antibiotics (first in 60 years) and then getting published in Nature.

https://www.nature.com/articles/d41586-023-03668-1

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#482
post #452

Earlier quoted context omitted.

"probably the single most important development in human history" is the kind of hyperbole you'd only find here. Better than medicine, agriculture, electrification, or music? That point of view simply does not jive with what I see so far from AI. It has had little impact beyond filling the internet with low-effort content. I feel like the crypto evangelists never got off the hype train. They just picked a new destina…

I mean maybe not the single most important development, but definitely a very important technological development with the potential to revolutionize multiple industries

Can I ask what industries with what application? I've seen lots of task like summarizing articles or producing text. The image and video work seems too rudimentary to be taken seriously.

Is there something out there that seems like a killer application?

I was amazed at the idea of the block chain but we never found a use for it outside of cryptocurrency. I see a similariy with AI hype.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#483
post #101

The way to view this kind of parasitism is how we look at patent trolls. When you look at the RIAA/MPAA lawsuits, while I don't agree with them, at least file sharing was basically a canonical form of copyright infringement. With LLMs we have an aspect of a text corpus that the creators were not using (the language patterns) and had no plans for or even idea that it could be used, and then when someone comes along an…

But it’s theirs, they created it and should therefore benefit from it. I’m honestly shocked at how much these companies are getting away with. It’s piracy on a massive scale. You can get a little discombobulated reading the comments from the nerds / subject idiots on this site.

It's "theirs"? The copyright monopoly was created to advance art and science, anything else is a mere perversion. So who are in the moral right, those advancing science by developing humanity's literal pinnacle of science, artficial intelligence, or those trying to hold the development back for their own commercial interest?

Mind you, Google books, literally just text from copyrighted books published for everyone online, was ruled "fair use", due to it's benefit to humanity.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#485
post #410
post #396

Earlier quoted context omitted.

If you're going to consider training ai as fair use, you'll have all kinds of different people with different skill levels training ais that work in different ways on the corpus. Not all of them will have the capability to cite a source, and plenty of them won't have it make sense to cite a source. Eg. Suppose I train a regression that guesses how many words will be in a book. Which book do I cite when I do an infere…

Any citation would be a good start. For complex subjects, I'm sure the citation page would be large, and a count would be displayed demonstrating the depth of the subject[3]. This is how Google did it with search results in the early days[1]. Most probable to least probable, in terms of the relevancy of the page. With a count of all possible results [2]. The same attempt should be made for citations.

Ok, now please cite the source of this comment you just made. It's okay if the citation list is large, just list your citations from most probably to the least probable.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#486
post #440

Earlier quoted context omitted.

Why isn't robots.txt enough to enforce copyright etc? If NYT didn't set robots.txt properly, is their content free-for-all? Yes I know the first answer you would jump to is "of course not, copyright is the default", but it's almost 2024 and we have had robots.txt as industry de jure to stop crawling.

robots.txt is not meant to be a mechanism of communicating the licensing of content on the page being crawled nor is it meant to communicate how the crawled content is allowed to be used by the crawler. Edit: same applies to humans. Just because a healthcare company puts up a S3 bucket with patient health data with “robots: *” doesn’t give you a right to view or use the crawled patient data. In fact, redistributing i…

Furthering the S3 health data thought exercise:

If OpenAI got their hands on an S3 bucket from Aetna (or any major insurer) with full and complete health records on every American, due to Aetna lacking security or leaking a S3 bucket, should OpenAI or any other LLM provider be allowed to use the data in its training even if they strip out patient names before feeding it into training?

The difference between this question or NYT articles is that this question asks about content we know should not be available publicly online (even though it is or was at some point in the past).

I guess this really gets at “do we care about how the training data was obtained or pre-processed, or do we only care about the output (a model’s weights and numbers, etc)

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#487

Earlier quoted context omitted.

It’s likely fair use.

> It's likely fair use. I agree. You can even listen to the NYT Hard Fork podcast (that I recommend btw https://www.nytimes.com/2023/11/03/podcasts/hard-fork-execut... ) where they recently had Harvard copyright law professor Rebecca Tushnet on as a guest. They asked her about the issue of copyrighted training data. Her response was: """ Google, for example, with the book project, doesn’t give you the full text and i…

Genuinely asking, is the “verbatim” thing set in stone? I mean, an entity spewing out NYTimes-like articles after having been trained on lots of NYTimes content sounds like a very grey zone, in the “spirit” of copyright law some may judge it as indeed not-lawful.

Of course, I’m not a lawyer and I know that in the US sticking to precedents (which mention the “verbatim” thing) takes a lot of precedence over judging something based on the spirit of the law, but stranger things have happened.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#488

Earlier quoted context omitted.

Warner Brothers sued and won against Asylum for this very thing lmao. https://en.m.wikipedia.org/wiki/Mockbuster

The Asylum suit was about trademark, not copyright. Asylum changed the title of the film to not infringe on Warner brother’s trademark and released it anyway.

yes, op asked about a video titled lotr.mkv, but with millions of tiny differences.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#489

Maybe the NYT should do more to protect its own content, it could go back to exclusively being a newspaper, they seem to understand that better than this whole internet funny business.

Yeah. But then how do you make money through all those ad impressions that you can ingest all over the internet right?
Post reply on HN