Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

801–810 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#801

I have deeply mixed feelings about the way LLMs slurp up copyrighted content and regurgitate it as something "new." As a software developer who has dabbled in machine learning, it is exciting to see the field progress. But I am also an author with a large catalog of writings, and my work has been captured by at least one LLM (according to a tool that can allegedly detect these things). Overall, current LLMs remind me…

> Overall, current LLMs remind me of those bottom-feeder websites that do no original research--those sites that just find an article they like, lazily rewrite it, introduce a few errors, then maybe paste some baloney "sources" (which always seems to disinclude the actual original source). That mode of operation tends to be technically legal, but it's parasitic and lazy and doesn't add much value to the world. Anothe…

That's the joke, these sites are long produced by LLMs. The result is obvious.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#802
post #767

Earlier quoted context omitted.

There are two problems with the “kid” analogy: a) In many closely comparable scenarios, yes, it’s copyright infringement. When Francis Ford Coppola made The Godfather film, he couldn’t just be “inspired” by Puzo’s book. If the story or characters or dialog are similar enough, he has to pay Puzo, even if the work he created was quite different and not a literal “copy”. b) Training an LLM isn’t like giving someone a bo…

> This copy is not a transitory copy in service of a fair use Training is almost certainly fair use, so it's exactly a transitory copy in service of fair use. Training, other than the brief "transitory copy" you mention is not copying, it's making a minuscule algorithmic adjustment based on fleeting exposure to the data.

Why is training “almost certainly” fair use?

Congress took the circuit holding in MAI Systems seriously enough to carve out a new fair use exception for copying software—entirely within the memory system of a licensed user—in service of debugging it.

If it took an act of Congress to make “unlicensed” debugging a fair use copy…

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#803
post #731

Earlier quoted context omitted.

Huh. I see downvotes. I am mystified, for if people and corporations are both treated stringently under the law, corporations will fight to have overly restrictive laws knocked down. I envision pitting corporate body against corporate body, when one corporatism lobbies, works to (for example) extend copyrights, others will work to weaken copyright. That doesn't happen as vigilantly currently, because there is no corp…

Corporations follow these laws much more stringently than individuals. Individuals often use pirated software to make things, I've seen many examples of that. I've never seen a corporation use pirated software to make things, they pay for licenses. Maybe there is some rare cases, but pirating is mostly a thing individuals do not corporations. So in general it is already as you say, corporations are much more targeted…

> I've also seen indie games use copyrighted material with no issues, but AAA titles seem to avoid that like the plague.

They use copyrighted material or they commit copyright infringement? The former doesn't necessarily constitute the latter. Likewise, given it's an option legally, there are other factors that go into the decision to use it that likely make it less attractive to AAA games.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#804

Earlier quoted context omitted.

I'm not for or against anything at this point until someone gets their balls out and clearly defines what copyright infringement means in this context. If you give a bunch of books to a kid all by the same author and then pay that kid to write a book in a similar style and then I go on to sell that book...have I somehow infringed copyright? The kids book at best is likely to be a very convincing facsimile of the orig…

There are two problems with the “kid” analogy: a) In many closely comparable scenarios, yes, it’s copyright infringement. When Francis Ford Coppola made The Godfather film, he couldn’t just be “inspired” by Puzo’s book. If the story or characters or dialog are similar enough, he has to pay Puzo, even if the work he created was quite different and not a literal “copy”. b) Training an LLM isn’t like giving someone a bo…

Regarding (b) ... while a specific method of training that involved persistent copying may indeed be a violation, it is far from clear that the general notion of "send server request for URL, digest response in software that is not a browser" is automatically a violation. If there is deemed to be a difference (i.e. all you are allowed to do without a license is have a human read it in a browser), then one can see training mechanisms changing to accomodate that.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#805

Earlier quoted context omitted.

There are two problems with the “kid” analogy: a) In many closely comparable scenarios, yes, it’s copyright infringement. When Francis Ford Coppola made The Godfather film, he couldn’t just be “inspired” by Puzo’s book. If the story or characters or dialog are similar enough, he has to pay Puzo, even if the work he created was quite different and not a literal “copy”. b) Training an LLM isn’t like giving someone a bo…

Regarding (b) ... while a specific method of training that involved persistent copying may indeed be a violation, it is far from clear that the general notion of "send server request for URL, digest response in software that is not a browser" is automatically a violation. If there is deemed to be a difference (i.e. all you are allowed to do without a license is have a human read it in a browser), then one can see tra…

It’s all about the purpose the transitory copy serves. The mechanism doesn’t really matter, so you can’t make categorical claims about (say) non-browser requests.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#806

Earlier quoted context omitted.

> I hope this results in Fair Use being expanded to cover AI training. Couldn't disagree more strongly, and I hope the outcome is the exact opposite. I think we've already started to see the severe negative consequences when the lion's share of the profits get sucked up by very, very few entities (e.g. we used to have tons of local papers and other entities that made money through advertising, now Google and Facebook…

Trying to prohibit this usage of information would not help prevent centralization of power and profit. All it would do is momentarily slow AI progress (which is fine), and allow OpenAI et al to pull the ladder up behind them (which fuels centralization of power and profit). By what mechanism do you think your desired outcome would prevent centralization of profit to the players who are already the largest?

> Trying to prohibit this usage of information

It's not trying to prohibit. If they want to use copyrighted material, they should have to pay for it like anyone else would.

> prevent centralization of profit to the players who are already the largest?

Having to destroy the infringing models altogether on top of retroactively compensating all infringed rightsholders would probably take the incumbents down a few pegs and level the playing field somewhat, albeit temporarily.

They'd have to learn how to run their business legally alongside everyone else, while saddled with dealing with an appropriately existential monetary debt.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#808
post #733
post #708

Earlier quoted context omitted.

“Feeling slighted” is a gross understatement of how a lack of compensation flowing to creators has shaped the internet and the wider world over the past 25 years. If we have a problem with the way top media companies compensate their creators, that is a separate issue - not a justification for layering another issue on top.

YouTube had made way more content creators wealthy than the NYT. Writers are not going to be paid more after this ruling either way.

Gadzooks! You're right! If only NYT had realised the secret to success was spewing out articles reacting to other articles reacting to other articles, they would all have been millionaires!

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#809

Earlier quoted context omitted.

How does AI compete with journalism? AI doesn't do investigative reporting, AI can't even observe the world or send out reporters. Which part of journalism is AI going to impact most? Opinion pieces that contain no new information? Summarizing past events?

AI certainly isn’t a replacement for journalism, but that doesn’t mean journalism will continue to exist if no one pays for it. If everyone gets their news from chatGPT or the like there will be no investigative reporting. We’re already beginning to see this with most people reading the google/Facebook blurbs instead of clicking the link and giving ad money let alone paying.

We're not just beginning to see it, it's already happened. It was enabled by digitized information, then amplified by the networking of the internet. The value of fresh information today is worth the price of a Google refresh, which for most people is effectively nothing. AI doesn't change that equation, and I'd argue it's overall impact on journalism will be less harmful than an ad-optimized economy or even the mere existence of YouTube.

Quality journalism hasn't had a meaningful source of funding for a while, now. If AI does end up replacing honest-to-goodness investigative reporting, it'll be for the same reason the internet replaced the newspaper.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#810

I have deeply mixed feelings about the way LLMs slurp up copyrighted content and regurgitate it as something "new." As a software developer who has dabbled in machine learning, it is exciting to see the field progress. But I am also an author with a large catalog of writings, and my work has been captured by at least one LLM (according to a tool that can allegedly detect these things). Overall, current LLMs remind me…

>All that aside, I tend to agree with the hypothesis that LLMs are a fad that will mostly pass. For professionals, it is really hard to get past hallucinations and the lack of citations. For writers maybe, but absolutely not for programmers, it's incredibly useful. I don't think anyone who's used GPT4 to improve their coding productivity would consider it a fad.

[deleted]
Post reply on HN