Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

251–260 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#251

Earlier quoted context omitted.

For some definitions of “working”.

Working enough that people and companies there exist, live, and are to some degree successful, yes. I've visited multiple times in the past few years and I found it to be pretty normal

My understanding is they have one of the most corrupt and unjust legal systems of the developed countries.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#253
When it comes to foundation models I think there needs to be a distinction between potential and actual infringement. You can use a broadly trained foundation model to generate copyright infringing content, just like you can use your brain to do so. But the fact that a model can generate such content doesn't mean it infringes by its mere existence. https://www.marble.onl/posts/general_technology_doesnt_viola...

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#254
post #101

The way to view this kind of parasitism is how we look at patent trolls. When you look at the RIAA/MPAA lawsuits, while I don't agree with them, at least file sharing was basically a canonical form of copyright infringement. With LLMs we have an aspect of a text corpus that the creators were not using (the language patterns) and had no plans for or even idea that it could be used, and then when someone comes along an…

But it’s theirs, they created it and should therefore benefit from it. I’m honestly shocked at how much these companies are getting away with. It’s piracy on a massive scale. You can get a little discombobulated reading the comments from the nerds / subject idiots on this site.

The copyright laws are unjust to begin with. Copyright shouldn’t last beyond ten years anyway. Simply claiming that piracy is always bad ignores the evil of the laws in the first place.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#255

Here come the innovation sponges. If this goes through then the models that the general public have access are going to be severely neutered while the ownership class will have a much better model that will never see the light of day due to legal risks and claims like this - therefore increasing the disparity between us all.

[flagged]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#256
The actual complaint is here:

https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20...

They state that "with minimal prompting", ChatGPT will "recite large portions" of some of their articles with only small changes.

I wonder why they don't sue the wayback machine first. You can get the whole article on the wayback machine. Not just portions. And not with small changes but verbatim. And you don't need any special prompting. As soon as you are confronted with a paywall window on the times websites, all you need to do is to go to the wayback machine, paste the url and you can read it.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#257

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

Trying to prevent AI from learning from copyrighted content would look completely stupid in a decade or two when we have AIs that are just as capable as humans, but solely due to being made of silicon rather than carbon are banned from reading any copyrighted material. Banning a synthetic brain from studying copyrighted content just because it could later recite some of that content is as stupid as banning a biologic…

It's not exactly a synthetic brain though, is it? LLMs are more like lookup tables for the texts they're trained on.

We will not have "AIs as capable as humans" in a couple decades. AIs will keep being tools used by humans. If you use copyrighted texts as input to a digital transformation, that's vopyright infringement. It's essentially the same situation as sampling in music, and imo the same solutions can be applied here: e.g. licenses with royalties.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#258
post #22
post #3

> The New York Times is suing OpenAI and Microsoft over claims the companies built their AI models by “copying and using millions” of the publication’s articles and now “directly compete” with the outlet’s content. Millions? Damn, they can churn out some content. 13 million[0]!. [0] https://archive.nytimes.com/www.nytimes.com/ref/membercenter... .

The NYT publishes about 200 pieces of journalism every day (according to their own website), and it was founded in 1851. That makes for a lot of articles.

The first 75+ years are no longer in copyright, so certainly possible to train on thousands maybe millions of NYT articles without concern.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#259
post #17

Earlier quoted context omitted.

Take estimated losses of the NYT from this "innovation" and multiply by 10^x where is "x" high enough to make tech companies stop and think before they break laws next time. That would be my approach at least.

which laws are broken exactly? it's not remotely settled law that "training an NN = copyright infringement"

[deleted]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#260

You do copyright for content that you invented and which didn't exist before. But NYT content is reporting on events truthfully to the public without any fiction or lies. Since there can be only one truth it should not matter whether NYT or Washington Post or ChatGPT is spinning it out. Unless NYT is claiming they don't report truth and publishes fiction. That is of concern since, NYT claims to reporth news truthfull…

Is your position that all non fiction textual work is uncopywritable?

Yeah. I don't think it makes much sense to allow copyright on textual descriptions of events that happened.

Now the question is whether did OpenAI violate the terms of service by using the bits transferred from NYT to train their LLM. I don't think their TOS had LLMs mentioned. So it's on NYT to be negligent and not update their TOS right?

Post reply on HN