Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

781–790 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#781

Will be interesting to see where this ends up. If I scrape the NYT content, and then commercialize a service that lets users query that content through an API (occasionally returning verbatim extracts) without any agreement from or payment to the NYT, that would be illegal. It's not obvious to me why putting an LLM in the middle of the process changes that.

As long as you pay for your copy of the content and the extracts are fair use, how would that be illegal?

It wouldn’t, but ‘fair use’ is doing a lot of work in that rhetorical. Seems like a court would be a proper place to define what is and isn’t.

Tbh I’m mostly curious about whether this settles out of court or whether it goes through the system and sets a precedent.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#782
post #412

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

I hope this results in OpenAI's code being released to everyone. This is way more important to humanity's future than any single software company. If OpenAI goes under, a dozen other outfits can replace them.

That'd be great!! I'd love it for their models to be open-sourced and replaced by a community effort, like WikiAI or whatever.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#783

Earlier quoted context omitted.

>It’s only fair for the copyright owners to want a share of the pie. No it's not, it's pure greed. Everyone'd think it absurd if copyright holders dared to demand that any human who reads their publicly available text has to pay them a fee, but just because OpenAI are training a brain made of silicon instead of a brain made of carbon all the rent-seekers come out to try to take advantage.

You know the NYT has to fork out money to build the content right ?

Do you really think they're losing subscribers to ChatGPT...? Is there a single real person that thinks, "Oh, I don't need to pay the NYT anymore, I can just wait for the next OpenAI update six months from now and it'll summarize all the news for me"?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#784
post #327

Earlier quoted context omitted.

Why can't AI at least cite its source? This feels like a broader problem, nothing specific to the NYTimes. Long term, if no one is given credit for their research, either the creators will start to wall off their content or not create at all. Both options would be sad. A humane attribution comment from the AI could go a long way - "I think I read something about this in the NYTimes on January 3rd, 2021." It appears t…

Why do you expect an AI to cite it's source? Humans are allowed to use and profit on knowledge they've learned from any and all sources without having to mention or even remember their sources. Yes, we all agree that it's better if they do remember and mention their sources, but we don't sue them for failing to do so.

Quite simply, if you're stating things authoritatively, then you should have a source.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#786
post #508

Earlier quoted context omitted.

Why isn't robots.txt enough to enforce copyright etc? If NYT didn't set robots.txt properly, is their content free-for-all? Yes I know the first answer you would jump to is "of course not, copyright is the default", but it's almost 2024 and we have had robots.txt as industry de jure to stop crawling.

Robot.txt isn't about copyrights, its about preventing bots. Its effectively a EULA. Copyright law only goes into effect when you distribute the content you scrape. If you scraped New York times for your own LLM that you used internally and didn't distribute the results, there would be no copyright infringement.

> If you scraped New York times for your own LLM that you used internally and didn't distribute the results, there would be no copyright infringement.

Why?

As far as I understand, the copyright owner has control of all copying, regardless of whether it is done internally or externally. Distributing it externally would be a more serious vilation, though.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#789
post #731

Earlier quoted context omitted.

Huh. I see downvotes. I am mystified, for if people and corporations are both treated stringently under the law, corporations will fight to have overly restrictive laws knocked down. I envision pitting corporate body against corporate body, when one corporatism lobbies, works to (for example) extend copyrights, others will work to weaken copyright. That doesn't happen as vigilantly currently, because there is no corp…

Corporations follow these laws much more stringently than individuals. Individuals often use pirated software to make things, I've seen many examples of that. I've never seen a corporation use pirated software to make things, they pay for licenses. Maybe there is some rare cases, but pirating is mostly a thing individuals do not corporations. So in general it is already as you say, corporations are much more targeted…

So then you refute the comment I replied to, and its parent.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#790
post #331

Earlier quoted context omitted.

For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…

I'm not for or against anything at this point until someone gets their balls out and clearly defines what copyright infringement means in this context. If you give a bunch of books to a kid all by the same author and then pay that kid to write a book in a similar style and then I go on to sell that book...have I somehow infringed copyright? The kids book at best is likely to be a very convincing facsimile of the orig…

Importantly, the kid- an individual human- got some wealth somewhat proportional to their effort. There’s non-trivial effort in recruiting the kid. We can’t clone the kid’s brain a million times and run it for pennies.

There are differences that are ethically, politically and in other ways between an AI doing something and a human doing the exact same thing. Those differences may need reflecting in new laws.

IANAL ans don’t have any positive suggestions for good laws, just pointing out that the analogy doesn’t quite hold. I think we’re in new territory where analogies to previous human activities aren’t always productive.

Post reply on HN