Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

621–630 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#621

Earlier quoted context omitted.

The rules we have now were made in the context of human brains doing the learning from copyrighted material, not machine learning models. The limitations on what most humans can memorize and reproduce verbatim are extraordinarily different from an LLM. I think it only makes sense to re-explore these topics from a legal point of view given we’ve introduced something totally new.

Human brains are still the main legal agents in play. LLMs are just a computer programs used by humans. Suppose I research for a book that I'm writing - it doesn't matter whether I type it on a Mac, PC, or typewriter. It doesn't matter if I use the internet or the library. It doesn't matter if I use an AI powered voice-to-text keyboard or an AI assistant. If I release a book that has a chapter which was blatantly cop…

I see two separate issues, the one you describe which is maybe slightly more clear cut: if a person uses an AI trained on copyrighted works as a tool to create and publish their own works, they are responsible if those resulting works infringe.

The other question, which I think is more topical to this lawsuit, is whether the company that trains and publishes the model itself is infringing, given they're making available something that is able to reproduce near-verbatim copyrighted works, even if they themselves have not directly asked the model to reproduce them.

I certainly don't have the answers, but I also don't think that simplistic arguments that the cat is already out of the bag or that AIs are analogous to humans learning from books are especially helpful, so I think it's valid and useful for these kinds of questions to be given careful legal consideration.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#622
post #570
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

> Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) Hacker News consistently have upvoted posts to let users circumvent paywalls. And even when it doesn't, conversations here (and on Twitter, Reddit, etc.) that summarize the articles and quote the relevant bi…

I don't think it's about scraping being a threat. It's that they violated the TOS and stand to make a ton of money from someone else's work.

I find irony in the newspaper suing AI when other news sources (admittedly not NYT) use AI to write the articles. How many other AI scrapers are just ingesting AI generated content?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#623

Earlier quoted context omitted.

The fact that it has never been done successfully outside performance arts.

That's a large category that includes everything from YouTubers to furry artists to live concerts.

Yes, but it's still a subset of the arts - it doesn't apply to movies, literature, nor even to the scripting for any of these.

And I should mention YouTubers wouldn't be making that much money if YouTube weren't enforcing copyright, as you could just upload their videos and get the ad money. Without copyright, you could also cut off their in-video promotions and add your own, including your own Patreon - so you would get 100% of the money off their work if you can out-promote them.

It's only live performances which are protected by the physical world's strict no-copying laws (the ones that don't allow the same macro object to be in two places at the same time).

So basically, no medium which allows copying of the works in whole or nearly whole has been successfully run with public works.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#624

I have deeply mixed feelings about the way LLMs slurp up copyrighted content and regurgitate it as something "new." As a software developer who has dabbled in machine learning, it is exciting to see the field progress. But I am also an author with a large catalog of writings, and my work has been captured by at least one LLM (according to a tool that can allegedly detect these things). Overall, current LLMs remind me…

>All that aside, I tend to agree with the hypothesis that LLMs are a fad that will mostly pass. For professionals, it is really hard to get past hallucinations and the lack of citations.

For writers maybe, but absolutely not for programmers, it's incredibly useful. I don't think anyone who's used GPT4 to improve their coding productivity would consider it a fad.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#625

I have deeply mixed feelings about the way LLMs slurp up copyrighted content and regurgitate it as something "new." As a software developer who has dabbled in machine learning, it is exciting to see the field progress. But I am also an author with a large catalog of writings, and my work has been captured by at least one LLM (according to a tool that can allegedly detect these things). Overall, current LLMs remind me…

Finally a reasonable take on this site.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#626

I have deeply mixed feelings about the way LLMs slurp up copyrighted content and regurgitate it as something "new." As a software developer who has dabbled in machine learning, it is exciting to see the field progress. But I am also an author with a large catalog of writings, and my work has been captured by at least one LLM (according to a tool that can allegedly detect these things). Overall, current LLMs remind me…

> Overall, current LLMs remind me of those bottom-feeder websites that do no original research--those sites that just find an article they like, lazily rewrite it, introduce a few errors, then maybe paste some baloney "sources" (which always seems to disinclude the actual original source). That mode of operation tends to be technically legal, but it's parasitic and lazy and doesn't add much value to the world.

Another way of looking at this is that bottom-feeder websites do work that could easily be done by an LLM. I've noticed a high correlation between "could be AI" and "is definitely a trashy click bait news source" (before LLMs were even a thing).

To be clear, if your writing could be replaced by an LLM today, you probably aren't a very good writer. And...I doubt this technology will stop improving, so I wouldn't make the mistake of thinking that 2023 will be a high point for LLMs and they aren't much better in 2033 (or whatever replaces them).

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#627

Earlier quoted context omitted.

They're not. They can skip the entirety of the NYT archives and not much of value will be lost. The issue is with every copycat lawsuit that sues every AI company out of existence. It's a chilling effect on AI development. Old entrenched companies trying to prohibit new ways of learning and sharing information for the sake of their profit.

Why don’t they train their AI on non-copyrighted material? It’s only fair for the copyright owners to want a share of the pie. I’d want one as well for my work.

Because they wouldn't have enough good quality training data then probably.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#628
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

> I do think they should quickly course correct at this point and accept the fact that they clearly owe something to the creators of content they are consuming.

Eventually these LLMs are going to be put in mechanical bodies with the ability to interact with the world and learn (update their weights) in realtime. Consider how absurd your perspective would be then, when it'd be illegal for this embodied LLM to read any copyrighted text, be it a book or a web page, without special permission from the copyright holder, while humans face no such restriction.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#629

I have deeply mixed feelings about the way LLMs slurp up copyrighted content and regurgitate it as something "new." As a software developer who has dabbled in machine learning, it is exciting to see the field progress. But I am also an author with a large catalog of writings, and my work has been captured by at least one LLM (according to a tool that can allegedly detect these things). Overall, current LLMs remind me…

LLMs are not a fad for many things especially programming. It improves my productivity at least by 100%. It’s also useful to understand specific and hard to Google questions or parsing docs quickly. I think it’s going to fizzle out for creative content though at least until these companies stop “aligning” it so much. Hard to be funny when you can’t even offend a single molecule.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#630

Earlier quoted context omitted.

He didn't circumvent computer security. He had had a right to use the MIT network and pull the JSTR information. He certainly did it in a shady way (computer in a closet) but it's every bit as arguable that he did it that way because he didn't want someone stealing or unplugging his laptop while it was downloading the data. He also did not distribute the information wholesale. What he planned on doing with the inform…

> right to use the MIT That right ended when he used it to break the law. It was also for use on MIT computers, not for remote access (which is why he decided to install the laptop, also knowing this was against his "right to use"). The "right to use" also included a warning that misuse could result in state and federal prosecutions. It was not some free for all. > and pull the JSTR information No, he did not have th…

Ok, so a lot you've written but it comes down to this. What law did he break?

Neither MIT nor JSTOR raised issue with what Schwartz did. JSTOR even went out of their way to tell the FBI they did not want him prosecuted.

Remember, again, with what he was charged. Wiretapping and intent to distribute. He wasn't charged with trespassing, breaking and entering, or anything else. Wiretapping and intent to distribute.

> His actions started to affect all users of JSTOR at MIT. The rate of outflow caused JSTOR to suffer performance, so JSTOR disabled all of MIT access.

And this is where you are confusing a "crime" with "misuse of a system". MIT and JSTOR were in their rights to cut access. That does not mean that what Schwartz did was illegal. Similar to how if a business owner tells you "you need to leave now" you aren't committing a crime because they asked you to leave. That doesn't happen until you are trespassed.

> Go ahead and demonstrate some wholesale distribution - pick an author and reproduce a few works, for example. I'll wait.

You violate copyright by transforming. And fortunately, it's really simple to show that chat GPT will violate and simply emit byte for byte chunks of copyrighted material.

You can, for example, ask it to implement Java's Array list and get several verbatim parts of the JDKs source code echoed back at you.

> How many could I get from what Schwartz downloaded?

0, because he didn't distribute.

Post reply on HN