Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

151–160 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#151

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

Computers aren't humans and LLMs aren't human brains.

We have no way to reconstruct memories from a preserved brain (yet). The exact ways in which humans form memories and store information isn't even known yet; we're still drilling into the specifics from higher-level concepts.

Modeling the human brain like nodes with weights ignores a lot of biological processes. Blood/oxygen flow, hormones, neurotransmitter decay, physical locality, chemical delays and interference from things like myelin sheaths, and other physical processes affect the synapses that are partially mirrored by computer simulations of neural networks. Unlike neural networks, human brains also don't work based on a single clock signal triggering input and output from all notes in instant steps.

Human memories are also not just "data in, weights out". They are heavily modified by things like mood, concentration, language(s) spoken, context, and emotional triggers. There's no way to feed a dictionary into a brain. Memory preservation consists of multiple stages, with differing memory types, involving various brain segments with dedicated functionality that can actually grow back due to neuroplasticity in some cases.

Efforts are being made to emulate living cells on computers, but LLMs aren't that. Inversely, efforts are also made to feed brain cells artificial signals and train them to play video games, which results in different behaviour compared to the systems we use for LLMs or other AI systems.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#152
post #141

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

It's quite different, is it not? I don't get the analogy. These models are scanning, storing, and ingesting more material than any one human could. Not only is the method completely different, the end goal and applications are as well. The analogy basically isn't one, at all. I'm pretty upset at companies using our personal data to make gobs of money off of. I'm also upset that they're now using our knowledge work to…

Consider a mega consulting firm with millions of von Neumann-like analysts. Together, they've processed the same vast data that LLMs have, but individually, none could. It's not an LLM, but its purpose is like ChatGPT: assisting clients with their tasks. If LLMs concern you due to their data processing, would a firm like this do the same?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#153

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

Humans don't have perfect recall, don't have virtually infinite storage and can't process requests in milliseconds. I wouldn't be surprised if the avenue of attack is that fair use laws are for humans, not robots, and if an AI has been trained on copyrighted data, that's not fair use. Also, don't forget that in reality what's happened is that a bunch of copyrighted text is encoded in the LLM in a way a human can't un…

[deleted]

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#154

Earlier quoted context omitted.

Humans don't have perfect recall, don't have virtually infinite storage and can't process requests in milliseconds. I wouldn't be surprised if the avenue of attack is that fair use laws are for humans, not robots, and if an AI has been trained on copyrighted data, that's not fair use. Also, don't forget that in reality what's happened is that a bunch of copyrighted text is encoded in the LLM in a way a human can't un…

Yup, and that reproducing it is a copyright violation. Still a copyright abolitionist though. Maybe now more people will join the fight?

Well if use the tool to reproduce copyrighted content you are violating copyright. But that’s not the primary usecase and nobody in their right mind is arguing that.

The weights are not a reproduction of the content. They are capable of it but so is a photocopier a lot more and we didn’t ban those either despite them technically being a lot more useful for violation.

Nah, this is expansionist doctrine and agenda for copyright - these companies are trying to copyright style and locations in latent space now.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#155
post #141

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

It's quite different, is it not? I don't get the analogy. These models are scanning, storing, and ingesting more material than any one human could. Not only is the method completely different, the end goal and applications are as well. The analogy basically isn't one, at all. I'm pretty upset at companies using our personal data to make gobs of money off of. I'm also upset that they're now using our knowledge work to…

The question I think is what is the scope of copyright?

I think historically it's about copying wholesale and redistributing for profit. That doesn't seem to be what's happening here.

If I read a publicly available article, and I create a summary, is that covered by copyright? If I get an AI to do that, is that somehow a 'special' type of summary that is covered?

Can a provider of content somehow say "you may not use this for summarization"? Or apply other terms to my consumption if they are making it publicly available?

I think the comparison is that if you ask a person something like:

"Describe a power station?"

They will likely have never been in a power station. They will be leaning heavily pieces of content they have consumed over the years, that presumably were copyrighted. If you create an article that describes a power station, have you breached copyright and should be sued by all the people over the years that have produced content about power stations?

Is it suddenly different when a computer does it?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#156

Copyrights and patents are holding back humanity.

Somebody posted this link in a comment on a thread the other day: https://news.artnet.com/market/koch-brother-loses-it-on-air-... And it occurred to me that this is precisely the thing that's holding back humanity: "Koch estimates that he has spent $25 million on legal fees—far more than the $5 million he originally spent on the fake wine itself." We have a legal system that is completely inaccessible to the average…

Two friends of mine, created a pretty good song some time back in 2013, and they tried to copyright it. One lawyer and mutual friend of mine and the band, asked for 200 euros to copyright just that one song. Later they realized, they could copyright the song for just 30 euros.

On some other totally unrelated news, my parents knew a person who sold en mass music on cassettes illegally copied from other cassettes back in the 80s. He bought a BMW and build a house just from that. I was friends with his grandson when we were teenagers, and his grandson didn't care to play basketball or football or anything like that. He was obsessed with listening to music and memorize all the lyrics and stuff.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#157

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

> Don't humans operate similarly?

I'm going to bypass this question a bit and say, who cares?

Why do we need to treat these things the same way we treat humans? Why can we not say that it's okay if a human does it, and not okay if it's a computer? There's nothing that requires us to establish 'fair' as treating them the same as people.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#158
post #141

Earlier quoted context omitted.

It's quite different, is it not? I don't get the analogy. These models are scanning, storing, and ingesting more material than any one human could. Not only is the method completely different, the end goal and applications are as well. The analogy basically isn't one, at all. I'm pretty upset at companies using our personal data to make gobs of money off of. I'm also upset that they're now using our knowledge work to…

Consider a mega consulting firm with millions of von Neumann-like analysts. Together, they've processed the same vast data that LLMs have, but individually, none could. It's not an LLM, but its purpose is like ChatGPT: assisting clients with their tasks. If LLMs concern you due to their data processing, would a firm like this do the same?

Bringing up poor fitting analogies won't change my opinion.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#159
IANAL, but copyright protections are pretty much tied to content and format and not to the idea itself, with the intent of preventing (or putting a price on) the copying of original works. The Times will have a very hard time proving that their content is being re-marketed by OpenAI. Having a competing product based on your ideas.

Compare:

"Steve Jobs [was] a tyrant": https://www.nytimes.com/2011/10/07/technology/steve-jobs-def...

Against:

"Whether to describe SJ as a tyrant is a matter of perspective...": https://chat.openai.com/share/28633f0c-007f-48b6-a615-1581c3...

The general way LLMs work do not preserve content in it's original form: the ideas they contain are extracted and clustered statistically - as a ELI5 refresher, an LLM reads 2 million NY Times articles and records that after the word "Steve" there are a lot of "Jobs" followed by a lot of "was a genius/tyrant", "founded Apple", etc. Then LLMs recreate the user question "Who was Steve Jobs?" using this complex net of token/word stats. Is that fair use? I think OpenAI lawyers will not even tap the fair use question, they will simply state that no copy happened, just a statistical collection of words from various sources.

And importantly: no LLM source is really prevalent, so the end result cannot be even be traced back to the source, especially if multiple, similar news sources are being fed to training. I have no idea how the Times is going to prove that its _theirs_ news.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#160
post #141

Earlier quoted context omitted.

It's quite different, is it not? I don't get the analogy. These models are scanning, storing, and ingesting more material than any one human could. Not only is the method completely different, the end goal and applications are as well. The analogy basically isn't one, at all. I'm pretty upset at companies using our personal data to make gobs of money off of. I'm also upset that they're now using our knowledge work to…

The question I think is what is the scope of copyright? I think historically it's about copying wholesale and redistributing for profit. That doesn't seem to be what's happening here. If I read a publicly available article, and I create a summary, is that covered by copyright? If I get an AI to do that, is that somehow a 'special' type of summary that is covered? Can a provider of content somehow say "you may not use…

> Is it suddenly different when a computer does it?

The answer is yes.

Everyone here knows where this is heading, and yet people will sit here and defend these companies as if they're on some righteous path. Humans are already becoming disposable statistics and the engines of compute, all for free and all for the benefit of corporations who didn't pay for any of it and don't even contribute back taxes.

Post reply on HN