Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

81–90 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#81
post #49

I don't think it's an exaggeration to say that LLMs might lead to the end of the open web, or at least a drastically reduced version of it. So much of these model's utility is in directly competing with the producers of the training data. Content creators and aggregators are seeing more and more reason to restrict and limit access, to avoid having AI companies consume all of their data and then be the ones making mon…

Is this open web in the room with us? Cheekiness aside, the myth of beautiful open web always seems to be exagerated. Most content on the web is already generated, ugly and spammy. Most of the traffic is already owned by mega corporations.

One could easily argue that it's unlikely it'll get worse - if anything, AI could empower competition as now a group of 3 passionate, free writers can compete with agenda-driven, for profit corporations on a similar level. This could very well make the web more open and free as it makes the web more accessible.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#82

Earlier quoted context omitted.

No it's the death of the corporate content hosting web - the open web was never about making money with your blog post/irc chat/usenet group/etc content, at least in my opinion. Let data be free! If someone wants to use it to make money, well, it's open, just like open source. It's still not okay to take open source work and claim it as your own, which is what copyright should be limited to. Stealing a photo or plagi…

Maybe I don’t want my non-corporate art to be a part of some large corporation’s training data.

Once again I have to point out that LLMs are not and will not be restricted to "large corporations". Even if initial training is prohibitively expensive, open base models already exist.

"corporate this" and "corporate that" are handy rhetorical flourishes but they distort the debate.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#83
post #49

I don't think it's an exaggeration to say that LLMs might lead to the end of the open web, or at least a drastically reduced version of it. So much of these model's utility is in directly competing with the producers of the training data. Content creators and aggregators are seeing more and more reason to restrict and limit access, to avoid having AI companies consume all of their data and then be the ones making mon…

Really? Seems like it's not so different conceptually to search engines. They also make money by indexing "training data" of a sort. But the open web trundles on regardless, because the search engines found ways to cut the content producers in on it. I see no reason why that can't be the case here too, with AI companies training their models to act more like search engines when data comes from certain sources - i.e.…

>Really? Seems like it's not so different conceptually to search engines. They also make money by indexing "training data" of a sort.

OK, so if a writer X has a blog to put up samples of their work to drive people to buy books and to get writing assignments and someone uses ChatGPT to write something in the style of X - this naively seems like a hit on that author's ability to sell their skills.

And I don't think it is fixable by making the AI act more like a search engine.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#84
post #49

I don't think it's an exaggeration to say that LLMs might lead to the end of the open web, or at least a drastically reduced version of it. So much of these model's utility is in directly competing with the producers of the training data. Content creators and aggregators are seeing more and more reason to restrict and limit access, to avoid having AI companies consume all of their data and then be the ones making mon…

Really? Seems like it's not so different conceptually to search engines. They also make money by indexing "training data" of a sort. But the open web trundles on regardless, because the search engines found ways to cut the content producers in on it. I see no reason why that can't be the case here too, with AI companies training their models to act more like search engines when data comes from certain sources - i.e.…

yes indeed, search engines settled that problem a long time ago in a civilized way. The more of these lawsuits, the harder the wild west of "training" will have to search for an acceptable solution to what is currently web-scale theft of data.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#86
post #60
post #36

Earlier quoted context omitted.

Do you think OpenAI doesn’t subscribe? I too load creative works into all sorts of temporary structures in order to read the paper. I don’t need to license it, I pay for a subscription. Should I pay more if I memorize the paper? Should I pay more if I read it to my sick friend in the hospital? Should I pay more if I save copies to my own hard drive and grep for words in the files? Should robots have a higher subscrip…

a subscription doesn't give you an automatic escape hatch out of copyright law. here's their ToS, which is pretty clear about what you cannot do: https://help.nytimes.com/hc/en-us/articles/115014893428-Term... (relevant parts below) Without NYT’s prior written consent, you shall not: ... (2) use robots, spiders, scripts, service, software or any manual or automatic device, tool, or process designed to data mine or sc…

actually if you had a subscription you might be more screwed than if you crawled it with free access, since having the subscription means agreement to the TOS.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#87
post #49

I don't think it's an exaggeration to say that LLMs might lead to the end of the open web, or at least a drastically reduced version of it. So much of these model's utility is in directly competing with the producers of the training data. Content creators and aggregators are seeing more and more reason to restrict and limit access, to avoid having AI companies consume all of their data and then be the ones making mon…

No it's the death of the corporate content hosting web - the open web was never about making money with your blog post/irc chat/usenet group/etc content, at least in my opinion. Let data be free! If someone wants to use it to make money, well, it's open, just like open source. It's still not okay to take open source work and claim it as your own, which is what copyright should be limited to. Stealing a photo or plagi…

if you want to give away your data for free, do it, but speak for yourself, not everyone is in position to be able to do it. Just like you can't force other people to give away their software (or anything else, really) for free if they choose to charge for it.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#88
post #71

Earlier quoted context omitted.

User-agent: GPTBot Disallow: /

ATMO you shouldn't have to maintain knowledge of what kind of crawler bot exist and having to maintain deny list. It should be the opposite, only expressedly allowed content should be crawled by mainaining allow lists.

You can do the opposite since the inception of robots.txt: User-agent: * Disallow: / and then whitelist google bot and whatnot. Most of the web is already configured this way. Just check robots.txt of any major website, e.g. https://twitter.com/robots.txt

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#89
post #72
post #9

Earlier quoted context omitted.

Paraphrasing is not the issue. The issue is that OpenAI copied the Times ’ creative works into a GPU to train a model. That copy was likely neither licensed nor fair use.

Google does the same to produce a search index.

you can easily opt out of that, or control it to your heart's desire (including what snippets to show), and it will be honored. There is no way to opt out of this bullshit. So no, not the same at all.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#90
> if a federal judge finds that OpenAI illegally copied The Times' articles to train its AI model, the court could order the company to destroy ChatGPT's dataset, forcing the company to recreate it using only work that it is authorized to use.

I'd like to see it happening but it sounds unrealistic.

Post reply on HN