Earlier quoted context omitted.
Just a question, do you remember a source for all the knowledge in your mind, or did you at least try to remember?
a computer isn't a human. aren't computers good at storing data? why can't they just store that data? they literally have sources in datasets. why can't they just reference those sources? human analogies are cute, but they're completely irrelevant. it doesn't change that it's specifically about computers, and doesn't change or excuse how computers work.
The New York Times is suing OpenAI and Microsoft for copyright infringement
461–470 of 912 posts
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#462Earlier quoted context omitted.
> AFAIK reading copyrighted works is not copyright infringement. [...] Are they trying to say that LLM training is a special type of reading that should be considered infringement? Nobody can argue that OpenAI was feeding the content to ChatGPT because ChatGPT was bored or was curious about current events. It was fed NYT's content so it would know how to reproduce similar content, for profit. I think getting a case-l…
> ... for profit. for non-profit
2. The non-profit OpenAI, Inc. company is not to be confused with the for-profit OpenAI GP, LLC [0] that it controls. OpenAI was solely a non-profit from 2015-2019, and, in 2019, the for-profit arm was created, prior to the launch of ChatGPT. Microsoft has a significant investment in the for-profit company, which is why they're included in this lawsuit.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#463Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#464AI will likely steamroll current copyright considerations. If we live in a world where anything can be generated at whim, copyright considerations will seem less and less relevant or even possible.
Wishful thinking, but maybe we'll all turn away from obsession with ownership, and instead turn to feeding the poor, clothing the naked, visiting the sick and afflicted.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#465Earlier quoted context omitted.
> I'm not sure if the verbatim content isn't more of a "stopped clock is right twice a day" or "monkeys typewriting shakespeare" situation. I think it’s more nuanced than that. Extending the “monkeys on typewriters” example, it would be like training and evolving those monkeys using Shakespeare as the training target. Eventually they will evolve to write content more Shakespeare like. If they get so close to the targ…
In the context of Shakespeare, I'd agree that there may be some competitive potential in the product. But in the context of news, something that evolves and relies on timely and accurate information, I don't see how something like that turns into competition for the NYT by being trained on past NYT outputs. If the argument is that people can use ChatGPT to get old NYT content for free, that can be illustrated simply…
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#466Earlier quoted context omitted.
The way I see it, if the NYT goes under (one of the biggest newspapers in the world), all similar outlets also go under. Major publishers, both of fiction and non-fiction, as well as images, video, and all other creative content, may also go under. Hence, there is no more (reliable) training data.
I'm not sure whether that would even be a net loss, TBH. So much commercial media is crap, maybe it would be better for the profit motive to be removed? On the fiction side, there's plenty of fan-fic and indie productions. On the nonfiction side, many indie creators produce better content these days than the big media outlets do. And there still might be room for premium investigative stories done either by a few con…
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#467The way to view this kind of parasitism is how we look at patent trolls. When you look at the RIAA/MPAA lawsuits, while I don't agree with them, at least file sharing was basically a canonical form of copyright infringement. With LLMs we have an aspect of a text corpus that the creators were not using (the language patterns) and had no plans for or even idea that it could be used, and then when someone comes along an…
But it’s theirs, they created it and should therefore benefit from it. I’m honestly shocked at how much these companies are getting away with. It’s piracy on a massive scale. You can get a little discombobulated reading the comments from the nerds / subject idiots on this site.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#468Earlier quoted context omitted.
Just a question, do you remember a source for all the knowledge in your mind, or did you at least try to remember?
a computer isn't a human. aren't computers good at storing data? why can't they just store that data? they literally have sources in datasets. why can't they just reference those sources? human analogies are cute, but they're completely irrelevant. it doesn't change that it's specifically about computers, and doesn't change or excuse how computers work.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#469Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#470Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…