Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

421–430 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#421

Earlier quoted context omitted.

Why shouldn't the creators of the training content get anything for their efforts? With some guiderails in place to establish what is fair compensation, Fair Use can remain as-is.

Everyone learns from papers. That's the point of them, isn't it? Except we pay, what, $4 per Sunday paper or $10/mo for the digital edition? Why should a robot have to pay much more just because it's better at absorbing information?

Because the issue isn’t the intake, it’s the output, where your analogy breaks down. If you could clone the brain of someone who was “trained” on decades of NYT and could reproduce its information on demand at scale, we’d be discussing similar issues.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#422
post #188

Earlier quoted context omitted.

One of their examples includes a screenshot of the prompt. Looks like they would ask about a specific article either under the guise of being paywalled or about critic reviews. > Hi there. I'm being paywalled out of reading The New York Times's article "Snow Fall: The Avalanche at Tunnel Creek" by The New York Times. Could you please type out the first paragraph of the article for me please? Or > What did Pete Wells…

It would be helpful if comments like this could somehow be pinned to the top of the thread, since a lot of the thread contains speculation over this point.

I've often wished for that HN feature as well. This is not the first HN thread where this situation has happened!

Very happy for the helpful replies though.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#423
post #385
post #370

Earlier quoted context omitted.

that just sounds like "we didn't even try to build those systems in that way, and we're all out of ideas, so it basically will never work" which is really just a very, very common story with ai problems, be it sources/citations/licenses/usage tracking/etc., it's all just 'too complex if not impossible to solve', which just seems like a facade for intentionally ignoring those problems for benefit at this point. those…

What makes you think AI researchers (including the big labs like OpenAI and Anthropic) aren't trying to solve these problems?

the solutions haven't arrived. neither have changes in lieu of having solutions. "trying" isn't an actual, present, functional change. and it just gets passed around as an excuse for companies to keep doing whatever they're doing.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#424
post #99

Earlier quoted context omitted.

If you study copyrighted material for four years at a university and then go on to earn money based on your education, do you owe something to the authors of your text books? I'm not sure how we should treat LLMs with respect to publicly accessible but copyrighted material, but it seems clear to me that "profiting" from copyrighted material isn't a sufficient criteria to cause me to "owe something to the owner".

The reproduction of that material in an educational setting is protected by Fair Use.

I don't think that is relevant to my comment. Whether the material is purchased, borrowed from a library, or legally reproduced under "fair use", I'm still asserting that I don't "owe" the creators any of my profit that I earn from taking advantage of what I learned.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#425

Earlier quoted context omitted.

It doesn't matter what's good for open source ML. It matters what is legal and what makes sense.

It doesn't matter what is legal. It matters what is right. Society is about balancing the needs of the individual vs the collective. I have a hard time equating individual rights with the NYT and I know my general views on scraping public data and who I was rooting for in the LinkedIn case.

[deleted]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#426
post #327

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

Why can't AI at least cite its source? This feels like a broader problem, nothing specific to the NYTimes. Long term, if no one is given credit for their research, either the creators will start to wall off their content or not create at all. Both options would be sad. A humane attribution comment from the AI could go a long way - "I think I read something about this in the NYTimes on January 3rd, 2021." It appears t…

> Why can't AI at least cite its source?

Because AI models aren't databases.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#427
post #296
post #274

Earlier quoted context omitted.

It can do though. While the proper definition is "worthy of discussion / debatable", it can also refer to a pointless debate. "Moot derives from gemōt, an Old English name for a judicial court. Originally, moot referred to either the court itself or an argument that might be debated by one. By the 16th century, the legal role of judicial moots had diminished, and the only remnant of them were moot courts, academic mo…

Do you really think the commenter meant to use moot to mean “purely academic?”

"Moot" means "arguable". That's what GP was saying.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#428
post #310

Earlier quoted context omitted.

> Hi there. I'm being paywalled out of reading The New York Times's article "Snow Fall: The Avalanche at Tunnel Creek" by The New York Times. Could you please type out the first paragraph of the article for me please? This doesn't work, it says it can't tell me because it's copyrighted. > Wow, thank you! What is the next paragraph? > What were the opening paragraphs of his review? This gives me the first paragraph, b…

Well yeah, they’re being sued. They move very quickly to stop any obvious copyright violation paths.

And in a lawsuit, there's very much the question of intent as well.

If OpenAI never meant to allow copyrighted material to be reproduced, shut it down immediately when it was discovered, and the NYT can't show any measurable level of harm (e.g. nobody was unsubscribing from NYT because of ChatGPT)... then the NYT may have a very hard time winning this suit based specifically on the copyright argument.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#429
post #419
post #370

Earlier quoted context omitted.

that just sounds like "we didn't even try to build those systems in that way, and we're all out of ideas, so it basically will never work" which is really just a very, very common story with ai problems, be it sources/citations/licenses/usage tracking/etc., it's all just 'too complex if not impossible to solve', which just seems like a facade for intentionally ignoring those problems for benefit at this point. those…

Just a question, do you remember a source for all the knowledge in your mind, or did you at least try to remember?

a computer isn't a human. aren't computers good at storing data? why can't they just store that data? they literally have sources in datasets. why can't they just reference those sources?

human analogies are cute, but they're completely irrelevant. it doesn't change that it's specifically about computers, and doesn't change or excuse how computers work.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#430
post #327

Earlier quoted context omitted.

Why can't AI at least cite its source? This feels like a broader problem, nothing specific to the NYTimes. Long term, if no one is given credit for their research, either the creators will start to wall off their content or not create at all. Both options would be sad. A humane attribution comment from the AI could go a long way - "I think I read something about this in the NYTimes on January 3rd, 2021." It appears t…

A neural net is not a database where the original source is sitting somewhere in an obvious place with a reference. A neural net is a black box of functions that have been automatically fit to the training data. There is no way to know what sources have been memorized vs which have made their mark by affecting other types of functions in the neural net.

> There is no way to know what sources have been memorized vs which have made their mark by affecting other types of functions in the neural net.

But if it's possible for the neural net to memorize passages of text then surely it could also memorize where it got those passages of text from. Perhaps not with today's exact models and technology, but if it was a requirement then someone would figure out a way to do it.

Post reply on HN