Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

431–440 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#431
post #293

Earlier quoted context omitted.

> Critically the question is, did the developers put reasonable guardrails in place to prevent it? Why? If I steal a bunch of unique works of art and store them in my house for only me to see, am I still committing a crime?

Yes, but policing affairs inside the home have always been impractical at the best of times. Of course, OpenAI and most other "AI" aren't affairs "inside the home"; they are affairs publicly demonstrated far and wide.

Not only not inside the home but also charging money for it.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#432
I don't see a judge ruling that training a model on copyrighted works to be infringement, I think (hope) that that is ruled to be protected as fair use. It's the LLM output behaviour, specifically the model's willingness to reproduce verbatim text which is clearly a violation of copyright, and should rightfully result in royalties being paid out. It also seems like something that should be technically feasible to filter out or cite, but with a serious cost (both in compute and in latency for the user). Verbatim text should be easy to identify, although it may require a Google Search - level amount of indexing and compute. As for summaries and text "in the style of" NYT or others, that's the tricky part. Not sure there's any high-precision way to identify that on the output side of an LLM, though I can imagine a GAN trained to do so (erring on the side of false-positives). Filtering-out suspiciously infringe-ish outputs and re-running inference seems much more solvable than perfect citations for non-verbatim output.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#433
post #99

Earlier quoted context omitted.

If you study copyrighted material for four years at a university and then go on to earn money based on your education, do you owe something to the authors of your text books? I'm not sure how we should treat LLMs with respect to publicly accessible but copyrighted material, but it seems clear to me that "profiting" from copyrighted material isn't a sufficient criteria to cause me to "owe something to the owner".

We don’t, and shouldn’t, give LLMs the same rights as people.

I think this is a misleading way to frame things. It is people who build, train, and operate the LLM. It isn't about giving "rights" to the LLM, it is about constructing a legal framework for the people who are creating LLMs and businesses around LLMs.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#434
post #203
post #109

Interesting. I think the appropriation, privatization, and monetization of "all human output" by a single (corporate) entity is at least shameless, probably wrong, and maybe outright disgraceful. But I think OpenAI (or another similar entity) will succeed via the Sackler defense - OpenAI has too many victims for litigation to be feasible for the courts, so the courts must preemptively decide not to bother with compen…

What concerns me, and I don’t see mentioned as much as I would expect, is: how will people be compensated for generating new content if ChatGPT takes over? I believe the innovation that will really “win” generative AI in the long term is one that figures out how to keep the model populated with fresh, relevant, quality information in a sustainable way. I think generative AI represents a chance to fundamentally rethin…

Agree. My view is we’re in the Napster moment and someone is going to invent the iTunes Music Store. Language models are a distribution mechanism for knowledge content—- in many cases more efficient and useful than the originally packaged materials (akin to how downloading a single pop song is greater than buying the album). It feels clear this is where we’re headed (verified, compensated content delivered through a new mechanism); this lawsuit is like the RIAA v. music sharing and the question is just if the current players in AI make it through or if someone else will come in and do iTunes.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#435
post #345

Earlier quoted context omitted.

It doesn't matter what's good for open source ML. It matters what is legal and what makes sense.

It matters what ends up being best for humanity, and I think there are cases to be made both ways on this

[deleted]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#436
post #229
post #189

Earlier quoted context omitted.

The ChatGPT subscription is more valuable because it's built on the theft of the NYT content and many other authors' work.

No. It is a technology that can massively accelerate human progress—I don’t buy the theory that OpenAI used NYTimes content that was not freely available to them. If you read all of the public internet you probably have lots of snippets of NYTimes articles. Regarding the reading of the wirecutter by a browser tool, I don’t know how much of it is available online without subscription (because I subscribe to the NYTime…

It may have been freely available to them, but that doesn’t mean that they are free to reproduce its contents or otherwise make use of it at scale.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#437

So the NYT wants its content to be indexed by search engines, therefore makes it public to all crawlers, but then complains when some crawlers use the content to train AI on it? This issue is about the NYT wanting to lure internet users into its biased and politically motivated news website (and make them pay for it). If the NYT wants to they can block crawlers and rely on loyal readers typing nytimes.com in the brow…

Yes, it’s ok to reproduce the headline and banner image for search visibility. Not the whole article.

Whether their coverage is biased or not is immaterial to their legal argument.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#438

The arguments about being able to mimic New York Times “style” are weak, but the fact that they got it to emit verbatim NY Times content seems bad for OpenAI: > As outlined in the lawsuit, the Times alleges OpenAI and Microsoft’s large language models (LLMs), which power ChatGPT and Copilot, “can generate output that recites Times content verbatim

Sarah Silverman is claiming the same thing about her book. But I've tried really hard to get ChatGPT to output sentences verbatim from her book and just can't get it to. In fact, I can't even get it to answer simple questions about facts that are in her book but nowhere else -- it just says it doesn't know. Similarly I haven't been able to reproduce any text in the NYT verbatim unless it's part of a common quote or p…

Maybe she needs to sue Goodreads too. It's most likely a way for her to claw relevance for her unmarketed book by attaching "AI" to it and also "poor artist" to her work.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#439

Earlier quoted context omitted.

Yep! I like having access to high-quality information and producing, collecting, editing, and publishing that is not free. Much of it is only cost-effective to produce if you can share it with a massive audience, I.e. sure if I want to read a great investigative piece on the corruption of a Supreme Court Justice I can hypothetically commission one, but in practice it seems much much better to allow people to have bus…

> that is not free Why did you specify that this stuff you like, you only like if it's "not free"? The hidden assumption is that the information you like wouldn't be made available unless someone was paying for it. But that's not in evidence; a lot of information and content is provided to the public due to other incentives: self-promotion, marketing, or just plain interest. Would you prefer not to have access to Wik…

I’ll restate it for clarity: I like high-quality information. Producing and publishing high-quality information is not free.

There are ways to make it free to the consumer, yes. One way is charity (Wikipedia) and another way is advertising. Neither is free to produce; the advertising incentive is also nuked by LLMs; and I’m not comfortable depending on charity for all of my information.

It is a lot cheaper to produce low-quality than high-quality information. This is doubly so in a world of LLMs.

There is ONE Wikipedia, and it is surely one of mankind’s crowning achievements. You’re pointing to that to say, “see look, it’s possible!”?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#440
post #331

Earlier quoted context omitted.

For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…

Why isn't robots.txt enough to enforce copyright etc? If NYT didn't set robots.txt properly, is their content free-for-all? Yes I know the first answer you would jump to is "of course not, copyright is the default", but it's almost 2024 and we have had robots.txt as industry de jure to stop crawling.

robots.txt is not meant to be a mechanism of communicating the licensing of content on the page being crawled nor is it meant to communicate how the crawled content is allowed to be used by the crawler.

Edit: same applies to humans. Just because a healthcare company puts up a S3 bucket with patient health data with “robots: *” doesn’t give you a right to view or use the crawled patient data. In fact, redistributing it may land you in significant legal trouble. Something being crawlable doesn’t provide elevated rights compared to something not crawlable.

Post reply on HN