Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

641–650 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#641
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

> the first being at the birth of modern search engines. Why do you say that? Search engines would at least direct the viewer to the source. NYT gets 35%+ of its traffic from Google: https://www.similarweb.com/website/nytimes.com/#traffic-sour...

That doesn’t mean that it wasn’t theft of their content. The internet would be a very different place if creator compensation and low friction micropayments were some of the first principles. Instead we’re left with ads as the only viable monetization model and clickbait/misinformation as a side effect.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#642
post #493

Earlier quoted context omitted.

By that logic you should have to pay the copyright holder of every library book you ever read, because you could later produce some content you memorised verbatim.

What do you actually believe, with that statement? Do you believe Libraries are operating illegally? That they aren't paying rightsholders? Also: GPT is not a legal entity in the united states. Humans have different rights than computer software. You are legally allowed to borrow books from the library. You are legally allowed to recite the content you read. You're not allowed to sell verbatim recitation of what you…

> Humans have different rights than computer software

Fortunately, the computer isn't the one being sued.

Instead it is the humans who use the computer. And those humans maintain their existing rights, even if they use a computer.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#643

I have deeply mixed feelings about the way LLMs slurp up copyrighted content and regurgitate it as something "new." As a software developer who has dabbled in machine learning, it is exciting to see the field progress. But I am also an author with a large catalog of writings, and my work has been captured by at least one LLM (according to a tool that can allegedly detect these things). Overall, current LLMs remind me…

I don’t view LLMs as a fad. It’s like drummers and drum machines. Machines and drummers co-exist really well. I think drum machines, among other things, made drummers better.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#644
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

The world you’re hoping for will put all AI tech only within the hands of the established top 10 media entities, who traditionally have never compensated fairly anyway.

Sorry but if that’s the alternative to some writers feeling slighted, I’ll choose for the writers to be sad and the tech to be free.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#645

Earlier quoted context omitted.

In your analogy, AI would be the videotape, not the person, because OpenAI is selling access to it.

I'm not so sure about that. It seems to me that they're selling me a service. Just like I might pay for a subscription to Adobe Photoshop or pay per-render fees to a rendering farm. I could use Photoshop to reproduce a copyrighted work, and in some circumstances (i.e. personal use) that'd be fine. Or I could use Photoshop to reproduce a copyrighted work and try to sell it for profit, which would clearly not be fine.…

The difference here is that Adobe is selling a set of tools that can recreate copyrighted work from the ground up. The Mona Lisa being previously incorporated into their tools is not a foundational necessity for their paintbrush to brush digital paint.

The same is not true for AI, which require copyrighted work be contained therein, in order for the tool part to function.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#646
post #634

Earlier quoted context omitted.

> I do think they should quickly course correct at this point and accept the fact that they clearly owe something to the creators of content they are consuming. Eventually these LLMs are going to be put in mechanical bodies with the ability to interact with the world and learn (update their weights) in realtime. Consider how absurd your perspective would be then, when it'd be illegal for this embodied LLM to read any…

A human faces the same restriction, if it provides commercial services on the internet creating code that is a copy of copyrighted code.

This isn't true; if you hire a contractor and tell them "write from memory the copyrighted code X which you saw before", and they have such a good memory that they manage to write it verbatim, then you take that code and use it in a way that breaches copyright, you're liable, not the person you paid to copy the code for you. They're only liable if they were under NDA for that code.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#647
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

> I do think they should quickly course correct at this point and accept the fact that they clearly owe something to the creators of content they are consuming. Eventually these LLMs are going to be put in mechanical bodies with the ability to interact with the world and learn (update their weights) in realtime. Consider how absurd your perspective would be then, when it'd be illegal for this embodied LLM to read any…

> while humans face no such restriction.

I have no idea what on earth you are talking about. People and corporations are sued for copyright infringement all the time.

https://copyrightalliance.org/copyright-cases-2022/

Reading and consuming other people content isn't illegal, but it also wouldn't be for a computer.

Reading and consuming content with the sole purpose of reproducing it verbatim is frowned upon, and can be sued, whether it's an LLM or a sweatshop in India.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#648
post #634

Earlier quoted context omitted.

A human faces the same restriction, if it provides commercial services on the internet creating code that is a copy of copyrighted code.

This isn't true; if you hire a contractor and tell them "write from memory the copyrighted code X which you saw before", and they have such a good memory that they manage to write it verbatim, then you take that code and use it in a way that breaches copyright, you're liable, not the person you paid to copy the code for you. They're only liable if they were under NDA for that code.

And what professional developer would not be under NDA for the code he produces for a corporation?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#649

I have deeply mixed feelings about the way LLMs slurp up copyrighted content and regurgitate it as something "new." As a software developer who has dabbled in machine learning, it is exciting to see the field progress. But I am also an author with a large catalog of writings, and my work has been captured by at least one LLM (according to a tool that can allegedly detect these things). Overall, current LLMs remind me…

> (according to a tool that can allegedly detect these things).

Eh, I would trust my own testing before trusting a tool that claims to have somehow automated this process without having access to the weights. Really it’s about how unique your content is and how similar (semantically) an output from the model is when prompted with the content’s premise.

I believe you, in any case. Just wanted to point out that lots of these tools are suspect.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#650

Earlier quoted context omitted.

> I do think they should quickly course correct at this point and accept the fact that they clearly owe something to the creators of content they are consuming. Eventually these LLMs are going to be put in mechanical bodies with the ability to interact with the world and learn (update their weights) in realtime. Consider how absurd your perspective would be then, when it'd be illegal for this embodied LLM to read any…

> while humans face no such restriction. I have no idea what on earth you are talking about. People and corporations are sued for copyright infringement all the time. https://copyrightalliance.org/copyright-cases-2022/ Reading and consuming other people content isn't illegal, but it also wouldn't be for a computer. Reading and consuming content with the sole purpose of reproducing it verbatim is frowned upon, and can…

>I have no idea what on earth you are talking about. People and corporations are sued for copyright infringement all the time.

They're sued for _producing content_, not consuming content. If a human takes copyrighted output from an LLM and publishes it, they're absolutely liable if they violated copyright.

>Reading and consuming other people content isn't illegal, but it also wouldn't be for a computer.

That is absolutely what people in this thread are suggesting should happen: that it should be illegal for OpenAI et. al. to train models on publicly available content without first receiving permission from the authors.

>Reading and consuming content with the sole purpose of reproducing it verbatim is frowned upon, and can be sued, whether it's an LLM or a sweatshop in India.

That's irrelevant here because people training LLMs aren't feeding them copyrighted content for the sole purpose of reproducing it verbatim.

Post reply on HN