Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

81–90 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#81

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

So Chinese LLMs are bad actors, but USA LLMs are the good guys?

I don't see it that way, but I'm sure from an American perspective that how it seems.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#82

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

This suggests to me that copyright laws are becoming out of date.

The original intent was to provide an incentive for human authors to publish work, but has become more out of touch since the internet allowed virtually free publishing and copying. I think with the dawn of LLMs, copyright law is now mainly incentivising lawyers.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#84
post #36

I think the train has left the station and the ship has sailed. I'm not sure it's possible to put this genie back in the bottle. I had stuff stolen by OpenAI too, and I felt bad about it (and even send them a nasty legal letter when it could output my creative work almost verbatim), but I think at this point, the legal landscape needs to somehow adjust. The Copyright Clause in the US Constitution is clear: To promote…

I see, the narrative switched form “cat’s out of the bag” to “genie’s out of the bottle”. Regardless, no one wants to ban llms. We just want the theft to stop.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#85

Earlier quoted context omitted.

Sarah Silverman is claiming the same thing about her book. But I've tried really hard to get ChatGPT to output sentences verbatim from her book and just can't get it to. In fact, I can't even get it to answer simple questions about facts that are in her book but nowhere else -- it just says it doesn't know. Similarly I haven't been able to reproduce any text in the NYT verbatim unless it's part of a common quote or p…

The complaint has specific examples they got from ChatGPT. There is a precedent: There were some exploit prompts that could be used to get ChatGPT to emit random training set data. It would emit repeated words or gibberish that then spontaneously converged on to snippets of training data. OpenAI quickly worked to patch those and, presumably, invested energy into preventing it from emitting verbatim training data. It…

1. The data emitted by that buffer-overflow-y prompt is both non-deterministic and actual training only appears a fraction of the time. There no prompt that allowed for reproducible targeting of data sets.

2. OpenAI's "patch" for that was to use their content moderation filter to flag those types of requests. They've done the same thing for copyrighted content requests. It's both annoying because those requests aren't against the ToS but it also shows that nothing has been inherently "fixed". I wouldn't even say it was patched.. they just put a big red sticker over it.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#86
post #72
post #39

Earlier quoted context omitted.

I would agree. Style is too amorphous (even among its own reporters and journalists, there are different styles), but verbatim repetition would be a problem. So what would the licensing be for all their content be (if presumably one could get ChatGPT to output all of the NYTs articles)? The unfortunate thing about these LLMs is they siphon all public data regardless of license. I agree with data owners one can’t Will…

FWIW When I was taking journalism classes, style was not amorphous. We had an entire book (400+ pages) which detailed every single specific stylistic rule we had to follow for our class. Had the same thing in high school newspaper. I can only assume that NYT has an internal one as well.

I wondered about that, but is that copyrightable? Can’t I use their style guide? If I did would the NYT sue me? If a writer who used it at the NYT went off on their own and started a substack and continued using the style, would they risk getting sued?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#87

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

They probably didn’t start with a lawsuit. They started asking for royalties. They probably didn’t get an offer they thought was fair and reasonable so they sued.

These media businesses have shareholders and employees to protect. They need to try and survive this technological shift. The internet destroyed their profitability but AI threatens to remove their value proposition.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#88
post #12

Can someone explain the technical difference between what search engines do to index newspapers versus what is being claimed here? Is the difference as simple as me being able to get summaries and content from a newspaper from GPT without needing to visit their website?

A search engines principle job is to provide you with links you can find the answer to your question. The LLMs are ingesting all of that content en masse and would provide you the answer directly, with no compensation to the writers who actually did the research to provide that answer. Search engines are symbiotic, LLMs are parasitic.

Except Google forced these companies to use their platform (Google's AMP) to host the content and essentially blackmailed into doing so ("we'll link directly, but only on page 3 of results").

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#89
You do copyright for content that you invented and which didn't exist before.

But NYT content is reporting on events truthfully to the public without any fiction or lies.

Since there can be only one truth it should not matter whether NYT or Washington Post or ChatGPT is spinning it out.

Unless NYT is claiming they don't report truth and publishes fiction.

That is of concern since, NYT claims to reporth news truthfully.

So is NYT scamming Americans hundreds of millions of dollars by charging for subscription fees by making a false promise on things that they report?

This should be the bigger question here.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#90

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

An LLM in Russia can commit the same crime in Russia, and get sued in Russia. No idea about China, but I know Russia has a working legal system.
Post reply on HN