Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

161–170 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#161

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

Access to ressources is hardly a new problem: when I was an NLP graduate student about a decade ago a teacher of us had scrapped (and continued to do so) a major newspaper for years to make a corpus. The legality of that was questionable at best, yet it was used in academic paper and a subset for training.

The same is equally applicable to image: Google got rich in part by making illegal copies of whatever image he could find. Existing regulations could be updated to include ML model but that won't stop bad or big enough actors to do what they want.

> We’re in a “mutually assured destruction” situation now

No, we aren't. Very good spam generators aren't comparable to mass destruction weapons.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#162
post #146

Earlier quoted context omitted.

AFAIK facts like happenings in the world are not copyrightable. So I guess the nyt is arguing it's copying their prose and way of writing about them?

Yes. Journalism is a job. People do the work of turning these happenings into words, and are paid for it. That's what's stolen here. The value created through doing that work. If it didn't have value, Microsoft would lose nothing by no longer ingesting it.

> People do the work of turning these happenings into words, and are paid for it. That's what's stolen here.

Stolen from whom? Journalists who got reported got paid. The owner is a billionaire. I don't understand your logic.

Does NYT pays money to the people/countries etc it uses to as subject to create content(NEWS)? Isn't that stealing then?

Also their website TOS didn't prohibit LLMs from using their data.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#163
post #36

I think the train has left the station and the ship has sailed. I'm not sure it's possible to put this genie back in the bottle. I had stuff stolen by OpenAI too, and I felt bad about it (and even send them a nasty legal letter when it could output my creative work almost verbatim), but I think at this point, the legal landscape needs to somehow adjust. The Copyright Clause in the US Constitution is clear: To promote…

I see, the narrative switched form “cat’s out of the bag” to “genie’s out of the bottle”. Regardless, no one wants to ban llms. We just want the theft to stop.

There is no theft. Hyperbole won't get you taken seriously, use correct terminology.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#164
post #123

Google can look up into their index and can remove whatever they want to, within minutes. But how that can be possible for an LLM? That is, "decontaminate" the model from certain parts of the corups? I can only think of excluding the data set from the training and then retrain? As a side note, I think LLM frenzy would be dead in few years, 10 years time frame at max. The rent seeking on these LLMs as of today would n…

> almost no workload (other than CAD, Graphics) runs on Windows or Unix including this very forum

About a fifth to a quarter of public-facing Web servers are Windows Server. Most famously, Stack Overflow[1].

[1]: https://meta.stackexchange.com/a/10370/1424704

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#165
post #21

For me it's quite obvious that if you make a profit from an engine that has as an input copyrighted material, then you owe something to the owner of this copyrighted content. We have seen this same problem with artists claiming stable diffusion engines were using their art.

Do all automakers that now develop electric cars owe Tesla something as they cashed in once they saw Tesla's successful copyrighted material l? A model is semantic, it contains the idea which is not copyrightable. Only how it is expressed could be copyrighted (i.e. if it outputs the copyright work verbatim). If this were not the case we would have plenty of monopolies and the world would fall apart.

https://www.tesla.com/blog/all-our-patent-are-belong-you

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#166
post #95
post #80

Earlier quoted context omitted.

Try selling subscriptions to your print-outs.

The equivalent analogy here is selling subscriptions to the printer, not the specific copyright infringing printout.

I hope HP isn't seeing this

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#167

Earlier quoted context omitted.

This argument is moot. Just because some countries - see china - steal intellectual property it doesnt mean we should. There are rules to the games we play specifically so we dont end up like them.

It's impossible to "steal" intellectual property without some kind of mind wiping device.

You must have used that device if you're making that argument in good faith.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#168
post #95
post #80

Earlier quoted context omitted.

Try selling subscriptions to your print-outs.

The equivalent analogy here is selling subscriptions to the printer, not the specific copyright infringing printout.

I disagree. A printer is too neutral - it's just a tool, like roads or the internet. Third parties can use them to commit copyright infringement, but that doesn't (or shouldn't) reflect on the seller of the tool.

I propose it's more like selling a music player that comes preloaded with (remixes of) recording artists' songs.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#169

If I were the CIA/US gov officials I would somehow want the NY Times to drop this case as one would not want AIs that don't have the talking points and propaganda pushed via papers not be part of the record. I am not saying that the NY Times is a CIA asset but from the crap they have printed in the past like the whole WMDs in Iraq saga and the puff piece of Elizabeth Holmes they are far from a completely independent…

> I am not saying that the NY Times is a CIA asset

Operation Mockingbird. While the publication as a whole may not be an asset, there are most assuredly assets within its staff.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#170
Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.)

I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if they put the horse before the cart. I do think they should quickly course correct at this point and accept the fact that they clearly owe something to the creators of content they are consuming. If they don’t, they are setting themselves up for a bigger loss down the road and leaving the door open for a more established competitor (Google) to do it the right way.

Post reply on HN