Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

471–480 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#471

Earlier quoted context omitted.

>the ship has sailed Certainly, but debating the spirit behind copyright or even "how to regulate AI" (a vast topic, to put it mildly) is only one possible route these lawsuits could take. I suspect that ultimately the winner is going to be business first (of course in the name of innovation), and the law second, and ethics coming last -- if Google can scan 129 million books [1] and store them without even a slap on…

The court decided to focus on the tiny snippets Google displayed rather than the full text on their servers backing the search functionality. The court found significant that Google deliberately limited the snippet view so it couldn't be used as a replacement for purchasing the original book. The opinion is a relatively easy read, I highly recommend it if you're interested in the issue. It's also notable the court co…

I referred to the case to point out that Google practically got away with what was a gigantic violation of the spirit or principle here, in that Google gets to keep a copy of these millions of works for itself without ever having paid for them, regardless of what it made available to the public.

As for "what would Google do with all these book copies anyway if they can't make it public?", that has now been answered more directly than ever.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#472
post #429
post #419

Earlier quoted context omitted.

Just a question, do you remember a source for all the knowledge in your mind, or did you at least try to remember?

a computer isn't a human. aren't computers good at storing data? why can't they just store that data? they literally have sources in datasets. why can't they just reference those sources? human analogies are cute, but they're completely irrelevant. it doesn't change that it's specifically about computers, and doesn't change or excuse how computers work.

When all the legal precedents we have are about humans, human analogies are incredibly relevant.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#473

Earlier quoted context omitted.

If NYT wins this, then there is going to be a massive push for payouts from basically everyone ever…I don’t see that wallet being fat for long.

If LLMs actually create added value and don't just burn VC money then they should be able to pay a fair price for the work of people they're relying upon. If your business is profitable only when you get your raw materials for free it's not a very good business.

Yup, and I think that'll quickly uncover the reality that LLMs do not generate enough value relative to their true cost. GPT+ already costs $20/month. M365 Copilot costs $30/user/month. They're already the most expensive B2B-ish software subscriptions out there, there's very little market room to add in more cost to cover payments to rightsholders.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#474
post #463

I think in the long run, it is in the interest of AI companies to incentivize creators to create high quality data! Not paying them their fair share will likely decrease the volume of high quality data available (or make it much less accessible). Unless these companies already have developed another architecture that can learn much more from the same dataset, the lack of new high quality data will be a problem for fu…

> Not paying them their fair share will likely decrease the volume of high quality data available

It won't. Thats not how capitalisam works. If high quality data became unavailable, then companies will be created to fix the problem. Only they look quite different from NYT.

Just like how Torrents didn't kill movie industry. These are lazy arguments made by people who want to make money through lawsuits.

Also I can guarentee you even in worst case, humanity would survive just fine without those high quality content just like it did for the past 50K+ years.

What you should actually be concerned about is stupid law suits like this that can prevent progress.

AI could help humanity solve more pressing problems like cancer.

By getting caught up in silly law suits like this and delaying progress one can make a case that you bring more suffering to the world.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#475
post #314

Earlier quoted context omitted.

Playing back large passages of verbatim content sold as your “product” without citation is almost certainly not fair use. Fair use would be saying “The New York Times said X” and then quoting a sentence with attribution. Thats not what OpenAI is being sued for. They’re being sued for passing off substantial bits of NYTimes content as their own IP and then charging for it saying it’s their own IP. This is also related…

> They’re being sued for passing off substantial bits of NYTimes content as their own IP and then charging for it saying it’s their own IP. In what sense are they claiming their generated contents as their own IP? https://www.zdnet.com/article/who-owns-the-code-if-chatgpts-... > OpenAI (the company behind ChatGPT) does not claim ownership of generated content. According to their terms of service, "OpenAI hereby assig…

They can’t transfer rights to the output of it isn’t theirs to begin with.

Saying they don’t claim the rights over their output while outputting large chunks verbatim is the old YouTube scheme of upload movie and say “no copyright intended”.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#476
post #411

Earlier quoted context omitted.

Presumably, if a passage of any significant length is cited verbatim (or almost verbatim), there would have been a way to track that source through the weights. The issue of replicating a style is probably more difficult.

> Presumably, if a passage of any significant length is cited verbatim (or almost verbatim), there would have been a way to track that source through the weights. Figure this out and you get to choose which AI lab you want to make seven figures at. It's a really difficult problem.

It’s likely first and foremost a resource problem. “How much different would the output be if that text hadn’t been part of the training data” can _in principle_ be answered by instead of training one model, training N models where N is the number of texts in the training data, omitting text i from the training data of model i, and then when using the model(s), run all N models in parallel and apply some distance metric on their outputs. In case of a verbatim quote, at least one of the models will stand out in that comparison, allowing to infer the source. The difficulty would be in finding a way to do something along those lines efficiently enough to be practical.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#477
post #314

Earlier quoted context omitted.

Playing back large passages of verbatim content sold as your “product” without citation is almost certainly not fair use. Fair use would be saying “The New York Times said X” and then quoting a sentence with attribution. Thats not what OpenAI is being sued for. They’re being sued for passing off substantial bits of NYTimes content as their own IP and then charging for it saying it’s their own IP. This is also related…

> They’re being sued for passing off substantial bits of NYTimes content as their own IP and then charging for it saying it’s their own IP. In what sense are they claiming their generated contents as their own IP? https://www.zdnet.com/article/who-owns-the-code-if-chatgpts-... > OpenAI (the company behind ChatGPT) does not claim ownership of generated content. According to their terms of service, "OpenAI hereby assig…

That part doesn't seem relevant to me in any case. IP pirates aren't prosecuted or sued because of a claim of ownership; they're prosecuted or sued over possession, distribution, or use.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#478
If meatspace’s non-technical industries and social/civil organs get the wool pulled over their eyes again, they deserve whatever tech-induced industry chaos that occurs - search/ad revenue models destroying journalism, social media destroying our civics and bonds, “move fast and break civil regs” (which have real people and their lives behind it) with Airbnb and rideshare, and now maybe LLMs and content ownership.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#479

Earlier quoted context omitted.

We all remember when Aaron Swartz got hit with a wire tapping and intent to distribute federal crime for downloading JSTR stuff right? It's really disgusting, IMO, that corporations that go above and beyond that sort of behavior are seeing NO federal investigations for this sort of behavior. Yet a private citizen does it and it's threats of life in prison. This isn't new, but it speaks to a major hole in our legal sy…

Circumventing computer security to copy items en masse to distribute wholesale without transformation is a far cry from reading data on public facing web pages.

He didn't circumvent computer security. He had had a right to use the MIT network and pull the JSTR information. He certainly did it in a shady way (computer in a closet) but it's every bit as arguable that he did it that way because he didn't want someone stealing or unplugging his laptop while it was downloading the data.

He also did not distribute the information wholesale. What he planned on doing with the information was never proven.

OpenAI IS distributing information they got wholesale from the internet without license to that information. Heck, they are selling the information they distribute.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#480
post #429
post #419

Earlier quoted context omitted.

Just a question, do you remember a source for all the knowledge in your mind, or did you at least try to remember?

a computer isn't a human. aren't computers good at storing data? why can't they just store that data? they literally have sources in datasets. why can't they just reference those sources? human analogies are cute, but they're completely irrelevant. it doesn't change that it's specifically about computers, and doesn't change or excuse how computers work.

LLMs are not databases. There is no "citation" associated with a specific query, any more than you can cite the source of the comment you just made.
Post reply on HN