Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

111–120 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#111
post #29

The challenge for all these AI companies is that the only thing of value for building a defensible commercial product is having proprietary datasets for training. With the underlying techniques and algorithms all being rapidly commoditized the power lies in who holds and owns that data. Like all other ML “revolutions” it’s the training data that matters and if one doesn’t have access to training data others don’t hav…

And I imagine that Gmail makes google very very special in this regard

FB likewise.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#112

The arguments about being able to mimic New York Times “style” are weak, but the fact that they got it to emit verbatim NY Times content seems bad for OpenAI: > As outlined in the lawsuit, the Times alleges OpenAI and Microsoft’s large language models (LLMs), which power ChatGPT and Copilot, “can generate output that recites Times content verbatim

I'm not sure if the verbatim content isn't more of a "stopped clock is right twice a day" or "monkeys typewriting shakespeare" situation. As I see it, most of the value in something like the NYT is as a trusted and curated source of information with at least some vetting. The content regurgitated from an LLM would be intermixed with false information and all sorts of other things, none of which are actually news from…

> I'm not sure if the verbatim content isn't more of a "stopped clock is right twice a day" or "monkeys typewriting shakespeare" situation.

I think it’s more nuanced than that.

Extending the “monkeys on typewriters” example, it would be like training and evolving those monkeys using Shakespeare as the training target.

Eventually they will evolve to write content more Shakespeare like. If they get so close to the target that some of them start reciting the Shakespeare they were trained on, you can’t really claim it was random.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#113

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

I don't think they're looking to prevent the inevitable, but rather see a target with a fat wallet from which a lot of money can be extracted. I'm not saying this in a negative way, but much of the "this is outrageous!" reaction to AI hasn't been about the building of models, but rather the realization that a few players are arguably getting very rich on those models so other people want their piece of the action.

If NYT wins this, then there is going to be a massive push for payouts from basically everyone ever…I don’t see that wallet being fat for long.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#114

Earlier quoted context omitted.

The second paragraph of the article is > As outlined in the lawsuit, the Times alleges OpenAI and Microsoft’s large language models (LLMs), which power ChatGPT and Copilot, “can generate output that recites Times content verbatim, closely summarizes it, and mimics its expressive style.” This “undermine[s] and damage[s]” the Times’ relationship with readers, the outlet alleges, while also depriving it of “subscription…

> closely summarizes it Absolutely not copyright infringement > mimics its expressive style Absolutely not copyright infringement > can generate output that recites Times content verbatim This one seems the closest to infringement, but still doesn't seem like infringement. A printer has this capability too. If a user told ChatGPT to recite NYT content and then sold that content, that would be 100% infringement, but w…

I'm in agreement, but this line is not quite an accurate metaphor:

> e.g. if someone printed out NYT articles and sold them, nobody would come after the printer manufacturer.

If the printer manufacturer had a product that could take one sentence and it would print multiple pages that complete a news article from that sentence, ...

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#115
post #12

Can someone explain the technical difference between what search engines do to index newspapers versus what is being claimed here? Is the difference as simple as me being able to get summaries and content from a newspaper from GPT without needing to visit their website?

The difference is that search engines don’t destroy the incentive to do the value-added activity of “produce original content.”

To the extent they do do that, e.g. Google’s “Knowledge Graph” snippets that extract content onto the results page, they also tend to be under fire for those. At least those (attempt to?) cite the source.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#116

You do copyright for content that you invented and which didn't exist before. But NYT content is reporting on events truthfully to the public without any fiction or lies. Since there can be only one truth it should not matter whether NYT or Washington Post or ChatGPT is spinning it out. Unless NYT is claiming they don't report truth and publishes fiction. That is of concern since, NYT claims to reporth news truthfull…

Not sure where you're coming from in this. A NYT article, once written is copyrighted. Using the content without attribution is at best plagiarism, and spitting it out the way the LLMs do is definitely a violation of if that copyright.

Unless you're telling me ChatGPT has eyes and sources just like the NYT and is worrying events as it sees them too?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#117
post #3

> The New York Times is suing OpenAI and Microsoft over claims the companies built their AI models by “copying and using millions” of the publication’s articles and now “directly compete” with the outlet’s content. Millions? Damn, they can churn out some content. 13 million[0]!. [0] https://archive.nytimes.com/www.nytimes.com/ref/membercenter... .

“Through Microsoft’s Bing Chat (recently rebranded as “Copilot”) and OpenAI’s ChatGPT, Defendants seek to free-ride on The Times’s massive investment in its journalism by using it to build substitutive products without permission or payment,” the lawsuit states. I can't be the only one that sees the irony of this news being "reported" and regurgitated over dozens of crappy blogs. ChatGPT [..] “can generate output tha…

Is your point that the NYT should sue bloggers? Or that given the existence of bloggers, they should not try to sue Microsoft? Or something else?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#118

You do copyright for content that you invented and which didn't exist before. But NYT content is reporting on events truthfully to the public without any fiction or lies. Since there can be only one truth it should not matter whether NYT or Washington Post or ChatGPT is spinning it out. Unless NYT is claiming they don't report truth and publishes fiction. That is of concern since, NYT claims to reporth news truthfull…

> You do copyright for content that you invented and which didn't exist before. I dont think that's accurate. The Copyright Act, § 103, allows copyright protection for "compilations (of facts)", as long as there is some "creative" or "original" act involved in developing the compilation, such as in the selection (deciding which facts to include or exclude) and arrangement (how facts are displayed and in what order).

Okay. But ChatGPT doesn't spin out the fact in the same order right? So how does this stand in court?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#119
post #3

> The New York Times is suing OpenAI and Microsoft over claims the companies built their AI models by “copying and using millions” of the publication’s articles and now “directly compete” with the outlet’s content. Millions? Damn, they can churn out some content. 13 million[0]!. [0] https://archive.nytimes.com/www.nytimes.com/ref/membercenter... .

“Through Microsoft’s Bing Chat (recently rebranded as “Copilot”) and OpenAI’s ChatGPT, Defendants seek to free-ride on The Times’s massive investment in its journalism by using it to build substitutive products without permission or payment,” the lawsuit states. I can't be the only one that sees the irony of this news being "reported" and regurgitated over dozens of crappy blogs. ChatGPT [..] “can generate output tha…

All those blogs are _also_ violating copyright, so I don't see the irony? One doesn't spend a million dollars suing a defendant with pennies to their name.

I'd also expect the Times style complaint to have merit because it's probably much easier for ChatGPT to imitate the NYT style than an arbitrary style.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#120

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

This argument is moot. Just because some countries - see china - steal intellectual property it doesnt mean we should. There are rules to the games we play specifically so we dont end up like them.

Ok, let’s address this from the standpoint of a node in the network of the thoughtscape. A denizen of the “inter”net, and also a victim of the exploitive nature of artists.

Media amalgamated power by farming the lives of “common” people for content, and attempt to use that content to manage lives of both the commons and unique, under the auspice of entertainmet. Which in and of itself is obviously a narrative convention which infers implied consent (id ask to what facetiously).

Keepsake of the gods if you will…

We are discussing these systems as though they are new (ai and the like, not the apple of iOS), they are not…

this is an obfuscation of the actual theft that’s been taking place (against us by us, not others).

There is something about reaping what you sow written down somewhere, just gotta find it.

-mic

Post reply on HN