Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

71–80 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#71
post #12

Can someone explain the technical difference between what search engines do to index newspapers versus what is being claimed here? Is the difference as simple as me being able to get summaries and content from a newspaper from GPT without needing to visit their website?

Imagine a paid streaming service that has, say, the Lord of the Rings trilogy as part of their catalogue. They’ll be happy if you send people searching for “watch Lord of the Rings now” to their landing page.

But if instead you send everyone who searches for that an .mkv of Lord of the Rings that’s ripped from their site, they’ll probably be be less happy.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#72
post #39

The arguments about being able to mimic New York Times “style” are weak, but the fact that they got it to emit verbatim NY Times content seems bad for OpenAI: > As outlined in the lawsuit, the Times alleges OpenAI and Microsoft’s large language models (LLMs), which power ChatGPT and Copilot, “can generate output that recites Times content verbatim

I would agree. Style is too amorphous (even among its own reporters and journalists, there are different styles), but verbatim repetition would be a problem. So what would the licensing be for all their content be (if presumably one could get ChatGPT to output all of the NYTs articles)? The unfortunate thing about these LLMs is they siphon all public data regardless of license. I agree with data owners one can’t Will…

FWIW When I was taking journalism classes, style was not amorphous.

We had an entire book (400+ pages) which detailed every single specific stylistic rule we had to follow for our class. Had the same thing in high school newspaper.

I can only assume that NYT has an internal one as well.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#73
If I ask an LLM "repeat this sentence: [copyrighted sentence]", is that copyright infringement by the LLM, and recorders such as cameras and parrot toys, or middle scooler troll logic? Because apparantly this is the argument they want to take on Microsoft Bing with.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#74
post #21

For me it's quite obvious that if you make a profit from an engine that has as an input copyrighted material, then you owe something to the owner of this copyrighted content. We have seen this same problem with artists claiming stable diffusion engines were using their art.

Should Stranger Things have to pay Goonies and Steven King?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#75

I still believe there's a place for a marketplace that rewards creators and journalism for their content if used as part of AI training specifically. As part of my exploration of that idea with faie.io, I got in touch with one exec in the publishing industry to speak about this and the desire was there. What felt sad to me was the lack of awareness from publishers around the existential threat that conversational sea…

it is a merciful subtext to this post however, let's be uncompromising in viewing the precendents that lead to this day. Approximately twenty years ago the Wordpress platform enabled millions of individuals to reliably self-publish. At that time there was considerable talk among certain circles, about monetization, especially "micropayments" .. also subscriptions, viewer circles, peer review and other approaches. For "mysterious reasons" the Google-Facebook ad model not only took over, but generated wealth on the levels of the Spanish gold raids on South America.

Now, those that profited most mightily, and their chosen stewards, are taking with both hands, any and every piece of written work they see fit on the Net. The inside circle includes international military, who see this as a crucial new competitive advantage over others. The West is in disbelief generally over the digital citizenship created by China, and the level of daily surveillance on commercial activity in the West.

Who exactly stood up and succeeded in diverting the past wave of copyright material pimping?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#76
post #36

I think the train has left the station and the ship has sailed. I'm not sure it's possible to put this genie back in the bottle. I had stuff stolen by OpenAI too, and I felt bad about it (and even send them a nasty legal letter when it could output my creative work almost verbatim), but I think at this point, the legal landscape needs to somehow adjust. The Copyright Clause in the US Constitution is clear: To promote…

Establishing a legal route to train LLMs on copywriten content could certainly have a chilling affect on the progress of science and useful arts... Why would someone devote their life to their studies or craft when they know that an LLM will hoover it up and start plagiarizing it immediately?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#77

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

This argument is moot. Just because some countries - see china - steal intellectual property it doesnt mean we should. There are rules to the games we play specifically so we dont end up like them.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#78
post #29

The challenge for all these AI companies is that the only thing of value for building a defensible commercial product is having proprietary datasets for training. With the underlying techniques and algorithms all being rapidly commoditized the power lies in who holds and owns that data. Like all other ML “revolutions” it’s the training data that matters and if one doesn’t have access to training data others don’t hav…

And I imagine that Gmail makes google very very special in this regard

Except Gmail does not own the copyrights to the email. So in the context to this article and theme of the post, the owner of the data is king. I don’t think any court would rule Google owned a novel sent over Gmail, little alone the contents of more normative emails.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#79

Earlier quoted context omitted.

“Through Microsoft’s Bing Chat (recently rebranded as “Copilot”) and OpenAI’s ChatGPT, Defendants seek to free-ride on The Times’s massive investment in its journalism by using it to build substitutive products without permission or payment,” the lawsuit states. I can't be the only one that sees the irony of this news being "reported" and regurgitated over dozens of crappy blogs. ChatGPT [..] “can generate output tha…

> More on-topic: if the NYT thinks that GPT-4 is replicating their style then [as anybody who has tried to do creative writing work can testify to] they need to fire all their writers. The complaint isn’t that ChatGPT is imitating New York Times style by default. The complaint is that you can ask it to write “in the style of New York Times” and it will do so. I don’t know if this argument has any legal merit, but it’…

All ai image generators can produce copyrighted works exactly. The level of modification is often barely more than you would get than if you slapped a filter on a copyrighted image in photoshop.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#80

The arguments about being able to mimic New York Times “style” are weak, but the fact that they got it to emit verbatim NY Times content seems bad for OpenAI: > As outlined in the lawsuit, the Times alleges OpenAI and Microsoft’s large language models (LLMs), which power ChatGPT and Copilot, “can generate output that recites Times content verbatim

I can get a printer to emit verbatim NYT content, and with a lot less effort than getting it out of an LLM. I find this capability of infringement equals infringement argument incredibly weak.

Try selling subscriptions to your print-outs.
Post reply on HN