Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

261–270 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#261
post #156

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

I see a complete economic collapse unless creators start getting paid both for their data upfront, and paid royalties when their data is used in an LLM response

Copyright doesn’t protect data, it only protects expression.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#262

I still believe there's a place for a marketplace that rewards creators and journalism for their content if used as part of AI training specifically. As part of my exploration of that idea with faie.io, I got in touch with one exec in the publishing industry to speak about this and the desire was there. What felt sad to me was the lack of awareness from publishers around the existential threat that conversational sea…

it is a merciful subtext to this post however, let's be uncompromising in viewing the precendents that lead to this day. Approximately twenty years ago the Wordpress platform enabled millions of individuals to reliably self-publish. At that time there was considerable talk among certain circles, about monetization, especially "micropayments" .. also subscriptions, viewer circles, peer review and other approaches. For…

Your comment covers lots of interesting topics. Talking only about chatGPT+Bing's threat to press and content creators survival, the difference with WordPress is that you could generate revenue "on your own" through your audience. Either with ads or through subscription. Conversational search will only make people visit less websites and consume the content directly from the AI. And sometimes without attribution to the original source, meaning your chances to monetize, as the creator/journalist/publisher, fall down to zero.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#263
post #2

NYT article with a lot more context https://www.nytimes.com/2023/12/27/business/media/new-york-t...

"Thank you. This article will be another labeled row in the next quarterly training batch. Now you can give our model a prompt to generate press coverage of any court case and it'll generate surpassing all the automated journalistic benchmarks that we have in place towards a safe, responsible, inclusive and aligned AI"

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#264
post #234

And here I am thinking it'd be amazing to have an AI that can on-demand read me every novel ever written. It'd be even cooler to jump into a text adventure game of any novel and have it actually follow the original text. I guess that clashes with our copyright world. (Is there hope of some kind of Netflix/Spotify model, with fractional royalties?)

I don't think anyone is saying that shouldn't be allowed. But there should be some model for consent/renumeration for the author of the original content.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#265

Earlier quoted context omitted.

If LLMs actually create added value and don't just burn VC money then they should be able to pay a fair price for the work of people they're relying upon. If your business is profitable only when you get your raw materials for free it's not a very good business.

By that logic you should have to pay the copyright holder of every library book you ever read, because you could later produce some content you memorised verbatim.

The difference here is scale. For someone to reproduce a book verbatim from memory it would take years of studying that book. For an LLM this would take seconds.

The LLM could reproduce the whole library quicker than a person could reproduce a single book.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#266

Earlier quoted context omitted.

What incentive do people have to publish work if their work is going to primarily be consumed by a LLM and spat out without attribution at people who are using the LLM?

Without 200 years of copyright protection, how will any author be able to afford food?

The fact that copyright protection is far too long is entirely separate from the need for some kind of copyright protection to exist at all. All evidence suggests that it's completely impossible to live off your work unless you copyright it for some reasonable period, with the possible exception of performance art (music, theater, ballet).

A writer or journalist just can't make money if any huge company can package their writing and market it without paying them a cent. This is not comparable to piracy, by the way, since huge companies don't move into piracy. But you try to compete with both Disney and Fox for selling your new script/movie, as an individual.

This experiment has also been tried to some extent in software: no company has been able to live off selling open source software. RedHat is the one that came closest, and they actually live by selling support for the free software they sell. Others like MySQL or Mongo lived by selling the non-GPL version of their software. And the GPL itself depends critically on copyright existing. Not to mention, software is still a best case scenario, since just having a binary version is often not enough, you need the original sources which are easy to guard even without copyright - no one cares so much for the "sources" of a movie or book.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#267
post #99
post #21

For me it's quite obvious that if you make a profit from an engine that has as an input copyrighted material, then you owe something to the owner of this copyrighted content. We have seen this same problem with artists claiming stable diffusion engines were using their art.

If you study copyrighted material for four years at a university and then go on to earn money based on your education, do you owe something to the authors of your text books? I'm not sure how we should treat LLMs with respect to publicly accessible but copyrighted material, but it seems clear to me that "profiting" from copyrighted material isn't a sufficient criteria to cause me to "owe something to the owner".

The reproduction of that material in an educational setting is protected by Fair Use.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#268

Earlier quoted context omitted.

And I imagine that Gmail makes google very very special in this regard

Except Gmail does not own the copyrights to the email. So in the context to this article and theme of the post, the owner of the data is king. I don’t think any court would rule Google owned a novel sent over Gmail, little alone the contents of more normative emails.

X might have a long-term edge if courts start ruling in favor of lawsuits like these? Being able to legally train on all Twitter data…

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#269

Earlier quoted context omitted.

On the other hand, you could also argue that if AI takes all financial incentives from professionals to produce original works, then the AI will lose out on quality material to train on and become worse. Unless your argument is there’s no need for anything else created by humanity, everything worth reading has already been written, and humanity has peaked and everyone should stop? Like all things, it’s about finding…

>financial incentives from professionals to produce original works People produce countless volumes of unpaid works of art and fiction purely for the joy of doing so; that's not going to change in future.

Anecdotal but I know lots of creatives (and by creatives I also include some devs) who've stopped publishing anything publicly because of various AI companies just stealing everything they can get their hands on.

They don't mind sharing their work for free to individuals or hell, to a large group of individuals and even companies, but AIs really take it to a whole different level in their eyes.

Whether this is a trend that will accelerate or even make a dent in the grand scheme of things, who knows, but at least in my circle of friends a lot of people are against AI companies (which is basically == M$) being able to get away with their shenanigans.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#270

Earlier quoted context omitted.

which laws are broken exactly? it's not remotely settled law that "training an NN = copyright infringement"

The legal argument, which I'm sure you are very well aware of, is that training a model on data, reorganizing, and then presenting that data as your own is copyright infringement.

I don't think OP is arguing in bad faith.The fact is it's unclear what laws this legal argument is supported by.
Post reply on HN