Live data from Hacker News

OpenAI: Copy, Steal, Paste

computerworld.com

41–50 of 79 posts

Re: OpenAI: Copy, Steal, Paste

#42
post #35

Can someone explain how adjusting weights based on viewing data is copyright infringement? I don't know if it's just not understood how these models work, or if it's just purposely misleading to try and cripple scary new tech. It seems like writers trying to complain it's copyright infringement to have other writers read their works for inspiration.

This argument, like everything else, is meaningless unless we talk about scale . It's like saying, well, I don't mind if someone sits by the road with pen and paper writing down license plate numbers. But now, when the scenario is adjusted to be a network of ALPN sensors blanketing an entire metro area, am I just "purposely misleading to try and cripple scary new tech" when I object to such a system?

I totally get that the scale is massive. But I am failing to see what copyright has to do with it.

I understand that AI is capable of copyright infringement, I too can draw a batman symbol off the top of my head, and it seems the fix for this simply output filtering AI content. But look at any AI-copyright complaint and all they focus on is training.

Re: OpenAI: Copy, Steal, Paste

#43
post #9

It seems obvious to me that a system in which intellectual property can be ingested into an artificial intelligence system and then provide no value to the original creator is unsustainable. The value of content that is generated through real-world effort like research and investigation will plummet towards zero. And then what? You can't just have a society built off of LLMs feeding each other. It also feels obvious…

The moment the rules are finalized, flocks of bright unprincipled people will rush to gamify them to turn a profit.

Re: OpenAI: Copy, Steal, Paste

#44
post #34
post #21

The (in my view) problem with the author's argument is that the first step he claims is happening, is not. Publicly available content gets read, as is the point of publicly publishing it. Then the user uses a computer program to make some statistics about the bit of content. Those bits of statistics about that specific work, on their own, cannot reproduce or recrate the specific work. Then those statistics are put in…

Forgive me for this one, but it comes from genuine curiosity and not snark. You are making assertions about how copyright law works, but you don't qualify it with either IANAL or any legal credentials. So I must ask: do you have a basis for these claims? I love participating in armchair analysis of the law, since in software we pretty much have no choice but to do so anyway, but my understanding has always been that…

> I think that from a legal standpoint, the actual technical means by which something was accomplished doesn't matter if the process as a whole is effectively copyright infringement.

Which is why when the user of the model prompts for something infringing, and is successful at getting close to verbatim output (because the prompt was too constraining, becuase the work is overrepresented in the training) it is that particular output that is infringing. And maybe that means that services operating that prompt/response software are guilty of contributory infringment if they can't adequetly prevent that kind of output.

But that doesn not mean that training the model was infringing. Nor does that mean distribution of the model is infringing. And if a user of the prompt/response software never prompts for anything infringing, and the software never spontaneously recreates anything infringing, there's no infringment happening.

There are lots of technologies out there that are highly capable of enabling infringment at a massive scale. And where the vast majority of their actual usage is absolutely infringing. But we don't completely shut down those technologies that on their own - are not infringing. Bittorrent clients are pefectly legal to develop. And distribute. And people use those clients to commit infringment at large scale. But they are still pefectly legal to write and distrubute.

Re: OpenAI: Copy, Steal, Paste

#45
post #21

The (in my view) problem with the author's argument is that the first step he claims is happening, is not. Publicly available content gets read, as is the point of publicly publishing it. Then the user uses a computer program to make some statistics about the bit of content. Those bits of statistics about that specific work, on their own, cannot reproduce or recrate the specific work. Then those statistics are put in…

There is a nice essay from 2004 that answers that question, "What Color Are Your Bits" (https://ansuz.sooke.bc.ca/entry/23, discussion https://news.ycombinator.com/item?id=24917679)

It talks about copyright infringement in music, but it applies just as well to AI training, just substitute "scrambled file" with "model weights":

> The scrambled file still has the copyright Colour because it came from the copyrighted input file. It doesn't matter that it looks like, or maybe even is bit-for-bit identical with, some other file that you could get from a random number generator. It happens that you didn't get it from a random number generator. You got it from copyrighted material; it is copyrighted. The randomly-generated file, even if bit-for-bit identical, would have a different Colour. The Colour inherits through all scrambling and descrambling operations and you're distributing a copyrighted work,

Re: OpenAI: Copy, Steal, Paste

#46
post #44
post #34

Earlier quoted context omitted.

Forgive me for this one, but it comes from genuine curiosity and not snark. You are making assertions about how copyright law works, but you don't qualify it with either IANAL or any legal credentials. So I must ask: do you have a basis for these claims? I love participating in armchair analysis of the law, since in software we pretty much have no choice but to do so anyway, but my understanding has always been that…

> I think that from a legal standpoint, the actual technical means by which something was accomplished doesn't matter if the process as a whole is effectively copyright infringement. Which is why when the user of the model prompts for something infringing, and is successful at getting close to verbatim output (because the prompt was too constraining, becuase the work is overrepresented in the training) it is that par…

TensorFlow is also perfectly legal to develop and distribute, and no one contests this.

People object to specific artifact, "model weights", which were produced using copyrighted works at the input, and can be used to reproduce those same copyrighted works back. In bittorrent analogy, people want to shut down specific pirate trackers and the pirate bay website.

Re: OpenAI: Copy, Steal, Paste

#47
I am not convinced that changing USA copyright law will do anything for NYT or protect any other content creators.

Once your content is online, it's subject to copy, piracy and manipulation of many forms, because it's on the WORLD WIDE WEB.

A US law can't enforce the world's network of computers and what they do with it's data.

Go ahead and sue OpenAI, I hope they lose.

I'm rooting for Open Source LLMs, that's where the real innovation is at, without censorship or copyright restrictions.

Re: OpenAI: Copy, Steal, Paste

#48
I'm a bit dubious about all these complaints. My snarky side wants to respond to

> By OpenAI's logic, any work you put online is fair game to be swiped and incorporated into the company's large language models.

With

> Any work you put online is fair game to be swiped and incorporated into a human's brain.

It's tough because the ability to copy an LLM is feasible, and an LLM is much more likely to be able to reproduce the underlying work with high fidelity, but there's no law that says that someone with a photographic memory can't look at your site data.

Re: OpenAI: Copy, Steal, Paste

#49

Technically it's not a copy. And nothing was stolen. It's a best fit curve among a series of datapoints. The datapoint is a copy, if the best fit curve never touches the datapoint then it's technically not a copy. The difficult part is the technicality here is legal and ethical from any standpoint. The high level ramifications are a bit unfair in the sense that yes the data is being used to create an AI that can repl…

Copyright laws do not support your argument. There have been many cases in music where the offending song was forced to pay because it was "close enough" to the curve but not touching it.

Because music has a lot of additional law written giving additional protections to song-writers independent of performers and recordings. That gives the abstract tonal sequence it's own copyright.

Re: OpenAI: Copy, Steal, Paste

#50
post #41

"Just. Pay. Me." Ok, how much? How would one determine all sources and their contribution weights for each produced completion from GPT?

Usually, you pay what they ask and, if you don't like the price, you don't use their work. It's very simple, but OpenAI skipping that step opens them up to courts deciding that price, if a price is actually owed. That price would probably be more than 0, which is what the authors have gotten so far.
Post reply on HN