Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

501–510 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#501
post #462

Earlier quoted context omitted.

> ... for profit. for non-profit

1. Non-profit != "not making a profit." A non-profit can still earn monetary profit, and many do. 2. The non-profit OpenAI, Inc. company is not to be confused with the for-profit OpenAI GP, LLC [0] that it controls. OpenAI was solely a non-profit from 2015-2019, and, in 2019, the for-profit arm was created, prior to the launch of ChatGPT. Microsoft has a significant investment in the for-profit company, which is why…

I know all that. But who did the training?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#502
post #490

When I was studying AI in grad school many years ago getting good big data sets was always an issue. It never occurred to me to just copy one without permission.

And therein lies the value of indexing huge amounts of data, which alphabet (google, youtube, etc.), Microsoft (bing, etc.), and similar companies have been doing for years now.

If it is legal to simply index a website, then why shouldn't it be legal to train a model in the very same data?

Of course, websites should have some option for declining data mining for ML/AI purposes, in the same way the can decline scraping/indexing in the robots.txt file.

But that ship has kind of sailed, unless the courts decide otherwise.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#503
post #476

Earlier quoted context omitted.

> Presumably, if a passage of any significant length is cited verbatim (or almost verbatim), there would have been a way to track that source through the weights. Figure this out and you get to choose which AI lab you want to make seven figures at. It's a really difficult problem.

It’s likely first and foremost a resource problem. “How much different would the output be if that text hadn’t been part of the training data” can _in principle_ be answered by instead of training one model, training N models where N is the number of texts in the training data, omitting text i from the training data of model i , and then when using the model(s), run all N models in parallel and apply some distance me…

each llm costs ($10-100) millions to train x billions of trainings data ~= $100 quadrillion dollars, so that is unofortunately out of reach of most countries.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#504
post #486
post #440

Earlier quoted context omitted.

robots.txt is not meant to be a mechanism of communicating the licensing of content on the page being crawled nor is it meant to communicate how the crawled content is allowed to be used by the crawler. Edit: same applies to humans. Just because a healthcare company puts up a S3 bucket with patient health data with “robots: *” doesn’t give you a right to view or use the crawled patient data. In fact, redistributing i…

Furthering the S3 health data thought exercise: If OpenAI got their hands on an S3 bucket from Aetna (or any major insurer) with full and complete health records on every American, due to Aetna lacking security or leaking a S3 bucket, should OpenAI or any other LLM provider be allowed to use the data in its training even if they strip out patient names before feeding it into training? The difference between this ques…

> should [they] be allowed to use this data in training…?

Unequivocally, yes.

LLMs have proved themselves to be useful, at times, very useful, sometimes invaluable assistants who work in different ways than us. If sticking health data into a training set for some other AI could create another class of AI which can augment humanity, great!! Patient privacy and the law can f*k off.

I’m all for the greater good.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#505
post #429
post #419

Earlier quoted context omitted.

Just a question, do you remember a source for all the knowledge in your mind, or did you at least try to remember?

a computer isn't a human. aren't computers good at storing data? why can't they just store that data? they literally have sources in datasets. why can't they just reference those sources? human analogies are cute, but they're completely irrelevant. it doesn't change that it's specifically about computers, and doesn't change or excuse how computers work.

OK, let's say you were given a source for an LLM output such as "Common Crawl/reddit/1000000 books collection". Would this be usefull? Probably not. Or do you want the chat system to operate magnitudes slower so it can search the peta bytes of sources and warn of similarities constantly for every sentence? That's obviously a huge waste of resources, it should probably be done by the users appropriately for their use case, such as these NY Times journalists which were easily able to find such similarities themselves for their use case of "specifically crafted prompts to output NY Times text".

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#506
post #419
post #370

Earlier quoted context omitted.

that just sounds like "we didn't even try to build those systems in that way, and we're all out of ideas, so it basically will never work" which is really just a very, very common story with ai problems, be it sources/citations/licenses/usage tracking/etc., it's all just 'too complex if not impossible to solve', which just seems like a facade for intentionally ignoring those problems for benefit at this point. those…

Just a question, do you remember a source for all the knowledge in your mind, or did you at least try to remember?

No, but I'm a human and treating computers like humans is a huge mistake that we shouldn't make.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#507

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

> if NYT goes under a dozen similar outlets can replace them overnight Not when there’s no money in journalism because the generative AIs immediately steal all content. If nyt goes under no one will be willing to start a news business as everyone will see it’s a money loser.

The NYT has been dying a slow death since long before ChatGPT came along.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#508
post #331

Earlier quoted context omitted.

For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…

Why isn't robots.txt enough to enforce copyright etc? If NYT didn't set robots.txt properly, is their content free-for-all? Yes I know the first answer you would jump to is "of course not, copyright is the default", but it's almost 2024 and we have had robots.txt as industry de jure to stop crawling.

Robot.txt isn't about copyrights, its about preventing bots. Its effectively a EULA. Copyright law only goes into effect when you distribute the content you scrape. If you scraped New York times for your own LLM that you used internally and didn't distribute the results, there would be no copyright infringement.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#509
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

It’s likely fair use.

It's likely not. Search for "the four factors of fair use". While I think OpenAI will have decent arguments for 3 of the factors, they'll get killed on the fourth factor, "the effect of the use on the potential market", which is what this lawsuit is really about.

If your "fair use" substantially negatively affects the market for the original source material, which I think is fairly clear in this case, the courts wont look favorably on that.

Of course, I think this is a great test case precisely because the power of "Internet scale" and generative AI is fundamentally different than our previous notions about why we wanted a "fair use exception" in the first place.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#510

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

> if NYT goes under a dozen similar outlets can replace them overnight Not when there’s no money in journalism because the generative AIs immediately steal all content. If nyt goes under no one will be willing to start a news business as everyone will see it’s a money loser.

How does AI compete with journalism? AI doesn't do investigative reporting, AI can't even observe the world or send out reporters.

Which part of journalism is AI going to impact most? Opinion pieces that contain no new information? Summarizing past events?

Post reply on HN