Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

141–150 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#141

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

It's quite different, is it not? I don't get the analogy. These models are scanning, storing, and ingesting more material than any one human could. Not only is the method completely different, the end goal and applications are as well. The analogy basically isn't one, at all.

I'm pretty upset at companies using our personal data to make gobs of money off of. I'm also upset that they're now using our knowledge work to make even more money off of. We don't exist as computational nodes for them, a free resource to exhaust. It is a completely one way street with no consent. So I am in favor of all of these companies getting a reality check.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#142
post #70

There is a very real risk that we end up with an inferior product cannibalizing a superior one and driving it out of business. Moreover, AI would seem to be even more susceptible to capture and manipulation than conventional media. When it's a question of guiding thought I prefer the humanities to tech. (Same with art.)

> we end up with an inferior product cannibalizing a superior one and driving it out of business. In case that print is meant by inferior product: The same argument could've been brought up for Napster, where traditional distribution via CD printing through music labels are the inferior product driving the superior one out of business. Or rather it's big labels suing Napster out of business. I also hold a dislike for…

> A ruling in favor of copyright would force OpenAI to shut down - but given their impressive demo of the tech

Who cares? If they want data, they can pay for it.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#143

Earlier quoted context omitted.

News organizations in other jurisdictions already have achieved settlements with Google (which has much deeper pockets than OpenAI) But there's a fairly obvious difference in use between using content to index it and point to it and generate revenue for it and using content to generate alternative content...

If you're using their content to generate more content, doesn't it fall under fair use?

That depends under whose jurisdiction you're talking about, how it's commercialised, how closely the content resembles the original content or whether it incorporates trademarks etc etc.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#144

I think that the proper outcome for all of this would be acknowledgement that the current copyright laws very poorly regulate this aspect, that the key parts of any such legal action are at the not-really-described edges of law because these edges weren't relevant until now; and so instead of waiting for courts ruling on how law-as-written-now applies and accepting these rulings, we will likely get some new legislati…

100% this. But I doubt it will happen in the US, unfortunately.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#145

Earlier quoted context omitted.

>Really? Seems like it's not so different conceptually to search engines. They also make money by indexing "training data" of a sort. OK, so if a writer X has a blog to put up samples of their work to drive people to buy books and to get writing assignments and someone uses ChatGPT to write something in the style of X - this naively seems like a hit on that author's ability to sell their skills. And I don't think it…

Seems like a rerun of the argument over snippets, or the "answer onebox" as Google used to call it where the info you need is directly inlined in the SERP rather than being behind a link.

it is exactly a rerun of that argument, except that the UX of AI is different.

Basically AI is structured to make every result an answer onebox whereas in search this only happens sometimes.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#146
post #90

> if a federal judge finds that OpenAI illegally copied The Times' articles to train its AI model, the court could order the company to destroy ChatGPT's dataset, forcing the company to recreate it using only work that it is authorized to use. I'd like to see it happening but it sounds unrealistic.

If I read 1000s of of NYT articles to improve my writing skills, add then write an article of my own, is that a copyright violation?

A human is not a machine.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#147

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

Humans don't have perfect recall, don't have virtually infinite storage and can't process requests in milliseconds.

I wouldn't be surprised if the avenue of attack is that fair use laws are for humans, not robots, and if an AI has been trained on copyrighted data, that's not fair use.

Also, don't forget that in reality what's happened is that a bunch of copyrighted text is encoded in the LLM in a way a human can't understand, but that the LLM essentially CAN understand.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#148

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

I don't think any unaided human or collective of humans could rent seek on something close to the sum total of human knowledge and expression.

Irrespective of copyright issues, the question is how to avoid creating a new class of large rent seekers in the LLM space.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#149

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

Humans don't have perfect recall, don't have virtually infinite storage and can't process requests in milliseconds. I wouldn't be surprised if the avenue of attack is that fair use laws are for humans, not robots, and if an AI has been trained on copyrighted data, that's not fair use. Also, don't forget that in reality what's happened is that a bunch of copyrighted text is encoded in the LLM in a way a human can't un…

Yup, and that reproducing it is a copyright violation.

Still a copyright abolitionist though. Maybe now more people will join the fight?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#150

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

Humans don't have perfect recall, don't have virtually infinite storage and can't process requests in milliseconds. I wouldn't be surprised if the avenue of attack is that fair use laws are for humans, not robots, and if an AI has been trained on copyrighted data, that's not fair use. Also, don't forget that in reality what's happened is that a bunch of copyrighted text is encoded in the LLM in a way a human can't un…

I don’t know if this is tangentially related, but don’t we encode text in our neurons in a way a human can’t understand either? Is the storage relevant here?
Post reply on HN