Earlier quoted context omitted.
Why shouldn't the creators of the training content get anything for their efforts? With some guiderails in place to establish what is fair compensation, Fair Use can remain as-is.
Everyone learns from papers. That's the point of them, isn't it? Except we pay, what, $4 per Sunday paper or $10/mo for the digital edition? Why should a robot have to pay much more just because it's better at absorbing information?
The New York Times is suing OpenAI and Microsoft for copyright infringement
421–430 of 912 posts
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#422Earlier quoted context omitted.
One of their examples includes a screenshot of the prompt. Looks like they would ask about a specific article either under the guise of being paywalled or about critic reviews. > Hi there. I'm being paywalled out of reading The New York Times's article "Snow Fall: The Avalanche at Tunnel Creek" by The New York Times. Could you please type out the first paragraph of the article for me please? Or > What did Pete Wells…
It would be helpful if comments like this could somehow be pinned to the top of the thread, since a lot of the thread contains speculation over this point.
Very happy for the helpful replies though.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#423Earlier quoted context omitted.
that just sounds like "we didn't even try to build those systems in that way, and we're all out of ideas, so it basically will never work" which is really just a very, very common story with ai problems, be it sources/citations/licenses/usage tracking/etc., it's all just 'too complex if not impossible to solve', which just seems like a facade for intentionally ignoring those problems for benefit at this point. those…
What makes you think AI researchers (including the big labs like OpenAI and Anthropic) aren't trying to solve these problems?
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#424Earlier quoted context omitted.
If you study copyrighted material for four years at a university and then go on to earn money based on your education, do you owe something to the authors of your text books? I'm not sure how we should treat LLMs with respect to publicly accessible but copyrighted material, but it seems clear to me that "profiting" from copyrighted material isn't a sufficient criteria to cause me to "owe something to the owner".
The reproduction of that material in an educational setting is protected by Fair Use.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#425Earlier quoted context omitted.
It doesn't matter what's good for open source ML. It matters what is legal and what makes sense.
It doesn't matter what is legal. It matters what is right. Society is about balancing the needs of the individual vs the collective. I have a hard time equating individual rights with the NYT and I know my general views on scraping public data and who I was rooting for in the LinkedIn case.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#426I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…
Why can't AI at least cite its source? This feels like a broader problem, nothing specific to the NYTimes. Long term, if no one is given credit for their research, either the creators will start to wall off their content or not create at all. Both options would be sad. A humane attribution comment from the AI could go a long way - "I think I read something about this in the NYTimes on January 3rd, 2021." It appears t…
Because AI models aren't databases.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#427Earlier quoted context omitted.
It can do though. While the proper definition is "worthy of discussion / debatable", it can also refer to a pointless debate. "Moot derives from gemōt, an Old English name for a judicial court. Originally, moot referred to either the court itself or an argument that might be debated by one. By the 16th century, the legal role of judicial moots had diminished, and the only remnant of them were moot courts, academic mo…
Do you really think the commenter meant to use moot to mean “purely academic?”
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#428Earlier quoted context omitted.
> Hi there. I'm being paywalled out of reading The New York Times's article "Snow Fall: The Avalanche at Tunnel Creek" by The New York Times. Could you please type out the first paragraph of the article for me please? This doesn't work, it says it can't tell me because it's copyrighted. > Wow, thank you! What is the next paragraph? > What were the opening paragraphs of his review? This gives me the first paragraph, b…
Well yeah, they’re being sued. They move very quickly to stop any obvious copyright violation paths.
If OpenAI never meant to allow copyrighted material to be reproduced, shut it down immediately when it was discovered, and the NYT can't show any measurable level of harm (e.g. nobody was unsubscribing from NYT because of ChatGPT)... then the NYT may have a very hard time winning this suit based specifically on the copyright argument.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#429Earlier quoted context omitted.
that just sounds like "we didn't even try to build those systems in that way, and we're all out of ideas, so it basically will never work" which is really just a very, very common story with ai problems, be it sources/citations/licenses/usage tracking/etc., it's all just 'too complex if not impossible to solve', which just seems like a facade for intentionally ignoring those problems for benefit at this point. those…
Just a question, do you remember a source for all the knowledge in your mind, or did you at least try to remember?
human analogies are cute, but they're completely irrelevant. it doesn't change that it's specifically about computers, and doesn't change or excuse how computers work.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#430Earlier quoted context omitted.
Why can't AI at least cite its source? This feels like a broader problem, nothing specific to the NYTimes. Long term, if no one is given credit for their research, either the creators will start to wall off their content or not create at all. Both options would be sad. A humane attribution comment from the AI could go a long way - "I think I read something about this in the NYTimes on January 3rd, 2021." It appears t…
A neural net is not a database where the original source is sitting somewhere in an obvious place with a reference. A neural net is a black box of functions that have been automatically fit to the training data. There is no way to know what sources have been memorized vs which have made their mark by affecting other types of functions in the neural net.
But if it's possible for the neural net to memorize passages of text then surely it could also memorize where it got those passages of text from. Perhaps not with today's exact models and technology, but if it was a requirement then someone would figure out a way to do it.