Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

211–220 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#211
post #109

Interesting. I think the appropriation, privatization, and monetization of "all human output" by a single (corporate) entity is at least shameless, probably wrong, and maybe outright disgraceful. But I think OpenAI (or another similar entity) will succeed via the Sackler defense - OpenAI has too many victims for litigation to be feasible for the courts, so the courts must preemptively decide not to bother with compen…

What does the Sackler defence refer to?

Opioids.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#212

Earlier quoted context omitted.

Do all automakers that now develop electric cars owe Tesla something as they cashed in once they saw Tesla's successful copyrighted material l? A model is semantic, it contains the idea which is not copyrightable. Only how it is expressed could be copyrighted (i.e. if it outputs the copyright work verbatim). If this were not the case we would have plenty of monopolies and the world would fall apart.

If Tesla thinks their competitors have violated any of their parents they are well within their rights to seek damages...

Ofcourse, my comment needs to be read in the context of what I'm responding to, they said input which I disagree with, output maybe there's a slight chance they have a case (depending on how openai has programmed it to output) and even then it's doubtful.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#213
post #101

The way to view this kind of parasitism is how we look at patent trolls. When you look at the RIAA/MPAA lawsuits, while I don't agree with them, at least file sharing was basically a canonical form of copyright infringement. With LLMs we have an aspect of a text corpus that the creators were not using (the language patterns) and had no plans for or even idea that it could be used, and then when someone comes along an…

But it’s theirs, they created it and should therefore benefit from it. I’m honestly shocked at how much these companies are getting away with. It’s piracy on a massive scale.

You can get a little discombobulated reading the comments from the nerds / subject idiots on this site.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#214

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

The NYT's strongest argument for infringement is that OpenAI is reproducing their content verbatim (and to make matters worse, without attribution). IANAL but it seems super likely to me that this will be found to be infringing sooner or later.

Do I really want to use a Chinese word processor that spits unattributed passages from the NYT into the articles I write? Once I publish that to my blog now I'm infringing and I can get sued too. Point is I don't see how output which complies with copyright law makes an LLM inferior.

The argument applies equally to code, if your use of ChatGPT, OpenAI etc. today is extensive enough, who knows what copyrighted material you may have incorporated illegally into your codebase? Ignorance is not a legal defense for infringement.

If anything it's a competitive advantage if someone develops a model which I can use without fear of infringement.

Edit: To me this all parallels Uber and AirBnB in a big way. OpenAI is just another big tech company that knew they were going to break the law on a massive scale, and said look this is disruptive and we want to be first to market, so we'll just do it and litigate the consequences. I don't think the situation is that exotic. Being giant lawbreakers has not put Uber or AirBnB out of business yet.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#215

Earlier quoted context omitted.

Establishing a legal route to train LLMs on copywriten content could certainly have a chilling affect on the progress of science and useful arts... Why would someone devote their life to their studies or craft when they know that an LLM will hoover it up and start plagiarizing it immediately?

The vast majority of quality art and is produced by people who do it because they want to create art, not for money, and most artists earn little.

Even if that is the case, plenty of art is made with the hope or dream that people will find it worth paying for, and there are many people out there who do in fact fully support themselves doing creative work. Having the copyright to that work is foundational to even be able to consider that possibility at all.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#216

Earlier quoted context omitted.

Oh thank goodness we can rely on charity for our information economy > Also, presumably NYT still has a business model unrelated to whatever OpenAI is doing with [NYT’s] data… That’s exactly the question. They are claiming it is destroying their business, which is pretty much self-evident given all the people in here defending the convenience of OpenAI’s product: they’re getting the fruits of NYTimes’ labor without p…

> Oh thank goodness we can rely on charity for our information economy You seem to be assuming an "information economy" should exist at all. Can you justify that?

Yep! I like having access to high-quality information and producing, collecting, editing, and publishing that is not free.

Much of it is only cost-effective to produce if you can share it with a massive audience, I.e. sure if I want to read a great investigative piece on the corruption of a Supreme Court Justice I can hypothetically commission one, but in practice it seems much much better to allow people to have businesses that undertake such matters and publish their findings to a large audience at a low unit price.

Now what’s your argument for removing such an incentive?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#217

You do copyright for content that you invented and which didn't exist before. But NYT content is reporting on events truthfully to the public without any fiction or lies. Since there can be only one truth it should not matter whether NYT or Washington Post or ChatGPT is spinning it out. Unless NYT is claiming they don't report truth and publishes fiction. That is of concern since, NYT claims to reporth news truthfull…

Is your position that all non fiction textual work is uncopywritable?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#218
post #74
post #21

For me it's quite obvious that if you make a profit from an engine that has as an input copyrighted material, then you owe something to the owner of this copyrighted content. We have seen this same problem with artists claiming stable diffusion engines were using their art.

Should Stranger Things have to pay Goonies and Steven King?

If they used copyrighted material or trademarks, they almost certainly _did_ pay the Goonies property rightholders and Stephen King for the privilege. Why would you think they didn't?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#219

Earlier quoted context omitted.

If LLMs actually create added value and don't just burn VC money then they should be able to pay a fair price for the work of people they're relying upon. If your business is profitable only when you get your raw materials for free it's not a very good business.

By that logic you should have to pay the copyright holder of every library book you ever read, because you could later produce some content you memorised verbatim.

> the copyright holder of every library book

gets paid

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#220

Earlier quoted context omitted.

This suggests to me that copyright laws are becoming out of date. The original intent was to provide an incentive for human authors to publish work, but has become more out of touch since the internet allowed virtually free publishing and copying. I think with the dawn of LLMs, copyright law is now mainly incentivising lawyers.

What incentive do people have to publish work if their work is going to primarily be consumed by a LLM and spat out without attribution at people who are using the LLM?

I notice this in myself, even though I've never particularly made money from published prose on the internet.

But (under different accounts) I used to be very active on both HN and reddit. I just don't want to be anymore now for LLM reasons. I still comment on HN, but more like every couple of weeks than every day. And I have made exactly one (1) comment on reddit in all of 2023.

I'm not the only one, and a lot of smaller reddit communities I used to be active on have basically been destroyed by either LLMs, or by API pricing meant to reflect the value of LLM training data.

Post reply on HN