Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

281–290 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#281

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

This suggests to me that copyright laws are becoming out of date. The original intent was to provide an incentive for human authors to publish work, but has become more out of touch since the internet allowed virtually free publishing and copying. I think with the dawn of LLMs, copyright law is now mainly incentivising lawyers.

The cost of copying and publishing has been almost irrelevant to the need for copyright at least since the times of the printing press. In fact, when copying books was extremely expensive work, copyright was not even that needed - the physical book was about as valuable as the contents, so no money was there to be made from copying someone else's work vs coming up with your own.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#282

Such incidents mark the end of an era. The diminishing relevance of traditional media in the digital age is afoot. I feel sorry for those who feed their families through this industry, but they need to learn and adapt before it's too late. Even if this lawsuit finds merit, it's akin to temporarily holding back a tsunami with a mere stick. A momentary reprieve, but not a sustainable solution. I agree with those who sa…

[deleted]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#283

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

> Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. Foreign companies can be barred from selling infringing products in the United States. Russian and Chinese consumers are less interested in English-language articles. I can’t really get beh…

> Russian and Chinese consumers are less interested in English-language articles.

Isn't it just one additional step to automatically translate them?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#284

Earlier quoted context omitted.

If this is inevitable (and I'm not saying it's not), who will produce high quality news content?

AI. And, I fear, it will be good.

Curious how AI gets the raw information if there are no reporters nor newspapers. Does AI go to meetings or interview politicians?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#285
Excellent! I am all for this type of contested reality with sources and derivative works. You feed in something you can’t claim it’s not being used in a way that isn’t allowed if you can’t explain how the fuck your little box works in the first place. I mean seriously pouring gasoline on yourself and playing with matches is about the same cause and effect of input output.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#286
post #273

Earlier quoted context omitted.

If LLMs actually create added value and don't just burn VC money then they should be able to pay a fair price for the work of people they're relying upon. If your business is profitable only when you get your raw materials for free it's not a very good business.

Imagine if tomorrow it was decided that every programmer had to pay out money for every single thing they went on the internet to learn about beyond official documentation, every Stack Overflow question they looked at, every question they went to a search engine to find. The amount of money was decided by a non-tech official who was in charge of figuring out how much of the money they earned was owed to the places th…

Except that every stackoverflow post is explicitly creative commons: https://stackoverflow.com/help/licensing

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#287
post #244

Earlier quoted context omitted.

I don't understand. So if New York times reported on a new laws of physics and put as an article will became copyrighted? Nobody would be able to talk about it and has to discover it by themselves? How is reporting on an event different from reporting on discovering a scientific law?

The exact words used to explain the scientific law are copyrighted by the writer (presumably the paper's authors). Rephrasings are not copywrited by the source, but by the rephrasing entity (e.g. the NYT, or a teacher that made a handout for their class). Copyright on scientific papers is most definitely a thing, by the way.

If the bar for copyright is as low as ordering of words, then I don't even know what to say.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#288
post #101

The way to view this kind of parasitism is how we look at patent trolls. When you look at the RIAA/MPAA lawsuits, while I don't agree with them, at least file sharing was basically a canonical form of copyright infringement. With LLMs we have an aspect of a text corpus that the creators were not using (the language patterns) and had no plans for or even idea that it could be used, and then when someone comes along an…

Amusing to see someone referring to anyone other than the megacorp controlled by fucking micro$oft hoovering as much data as they can, legally and otherwise, as a parasite.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#289

Such incidents mark the end of an era. The diminishing relevance of traditional media in the digital age is afoot. I feel sorry for those who feed their families through this industry, but they need to learn and adapt before it's too late. Even if this lawsuit finds merit, it's akin to temporarily holding back a tsunami with a mere stick. A momentary reprieve, but not a sustainable solution. I agree with those who sa…

I agree that the tide is turning, but I don't think the argument that actual criminal behavior (I don't know if that's what OpenAI did, but that's what NYT alleges) should be glossed over in the name of progress.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#290
If AI companies wanted to train their models on good content, they had a chance to create a second Renaissance. Funding artist collectives to create content for their models. Paying royalties to authors. Generally increasing the value of human art while creating a new form of expression.

Instead they do what every large corporation does and treat art like content. They are making loads of money off the backs of artists who are already underpaid and often undervalued and they didn't have the decency to ask for permission.

I know publishers don't treat authors much better. But I see this as NYT fighting for their journalists.

Post reply on HN