Earlier quoted context omitted.
> Just learn to recognize and punish plagiarism via RLHF OpenAI has created a $100bn company on this transfer. The Times may have an interest in a material fraction of that wealth.
The NYT is also worth a tiny fraction of that. If it looks like they might get anywhere, it might be better for OpenAI to buy them
NY Times copyright suit wants OpenAI to delete all GPT instances
251–260 of 921 posts
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#252Companies that have content all see dollar signs. NYT won't mind if you use their content to train LLMs - as long as they get a commission. Reddit will shut down their free API and make you pay to get training content. Discord is going to be selling content for AI training too - if they haven't already done so. Twitter is doing it. They didn't care before because LLMs were just experiments. Now we're talking trillion…
NYT do not "have" content, they create content. It's their raison d'etre.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#253We developers like to pretend that LLM's are akin to humans and that they've been using things like NYTimes like humans as educational material. But they are not. It's much simpler, proprietary writing is now integrated into the source code of OpenAI, it would be as if I would copy parts of other propriety code and copy paste it into my own codebase. Claiming copy paste is a natural evolving process of millions of ye…
> it would be as if I would copy parts of other propriety code and copy paste it into my own codebase. It's not copy-pasted; it's compressed in a lossy manner. Even GPT4 has nowhere near enough memory to store the entirety of its training data in a non-lossy compression format. Just likes how humans compress the information we read.
Humans don’t have the scale machines have and moreover humans aren’t sevices, that argument doesn’t fly.
I really think NYTs data isn’t that important and nor crucial, LLMs could’ve just elided it. However, it’s more about training on copyrighted data in general which is kind of crucial for OpenAi, they trained their LLMs indiscriminately on copyrighted content without any plan to share any profits.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#254It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
To me, there is a sense that the news, which is real information about the society that we currently live in, should be availabe to all participants of that society. The notion of being a good citizen requires that one stays informed. Books, movies, videogames etc. don't have that role and are more consumption goods.
How would you pay for news otherwise?
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#255It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
If the news/magazine doesn't want this they can simple serve a cut down or zero length article to all non-paying viewers! But they want that SEO, and they want that marketing.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#256Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#257Earlier quoted context omitted.
>or produced replicas from memory and sold them (the most comparable example), I’d be sued. This is not the most comparable example, because it's not what ChatGPT is doing. The most comparable example is if you were hired as a contractor and the employer asked you to write verbatim some copyright content you'd memorised. If the employer then published it, they'd be the one liable, not you. >Humans produce derivative…
> The most comparable example is if you were hired as a contractor and the employer asked you to write verbatim some copyright content you'd memorised. If the employer then published it, they'd be the one liable, not you. No, you'd both be liable. You are not allowed to create copies of a copyrighted work, even from memory, for any commercial purpose. Making it public or not is irrelevant. This is more obvious with s…
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#258It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
If NYT was a HN startup the link to the archived version would be banned and dang would be slamming the ban hammer.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#259It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
To me, there is a sense that the news, which is real information about the society that we currently live in, should be availabe to all participants of that society. The notion of being a good citizen requires that one stays informed. Books, movies, videogames etc. don't have that role and are more consumption goods.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#260It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
It is am ethical grey area, but if the paywall applied to all user agents, which would make it similar to say buying a Kindle book, then you might see that as pirating, whereas if you use an archive service that was served the HTTP response and cached it, then you are using a proxy UA. If the news/magazine doesn't want this they can simple serve a cut down or zero length article to all non-paying viewers! But they wa…