Live data from Hacker News

OpenAI: Copy, Steal, Paste

computerworld.com

31–40 of 79 posts

Re: OpenAI: Copy, Steal, Paste

#31
Can someone explain how adjusting weights based on viewing data is copyright infringement?

I don't know if it's just not understood how these models work, or if it's just purposely misleading to try and cripple scary new tech.

It seems like writers trying to complain it's copyright infringement to have other writers read their works for inspiration.

Re: OpenAI: Copy, Steal, Paste

#32
post #30
post #23

Earlier quoted context omitted.

> These AI systems have been built on the back of free labor for decades. Do you repay publishers for information you’ve summarized or new insights you’ve gained after consuming their content? How is it different when AI does the same thing?

Yes. They get a subscription fee or are paid by advertising. No one pays to advertise to an AI.

There are people who still don't block ads?

Re: OpenAI: Copy, Steal, Paste

#33
post #21

The (in my view) problem with the author's argument is that the first step he claims is happening, is not. Publicly available content gets read, as is the point of publicly publishing it. Then the user uses a computer program to make some statistics about the bit of content. Those bits of statistics about that specific work, on their own, cannot reproduce or recrate the specific work. Then those statistics are put in…

Why should a crime depend on how the tool used for the crime was created. Like 2 guys write 2 scripts that output a copyrighted poem, then the cool one guy, call him Sam can go free because he used algorithm A but second guy goes to jail because he used algorithm B where the crime is exactly identical.

Re: OpenAI: Copy, Steal, Paste

#34
post #21

The (in my view) problem with the author's argument is that the first step he claims is happening, is not. Publicly available content gets read, as is the point of publicly publishing it. Then the user uses a computer program to make some statistics about the bit of content. Those bits of statistics about that specific work, on their own, cannot reproduce or recrate the specific work. Then those statistics are put in…

Forgive me for this one, but it comes from genuine curiosity and not snark. You are making assertions about how copyright law works, but you don't qualify it with either IANAL or any legal credentials. So I must ask: do you have a basis for these claims?

I love participating in armchair analysis of the law, since in software we pretty much have no choice but to do so anyway, but my understanding has always been that we still don't actually have strong case-law for machine learning and AI. It does seem like the existing cases regarding weights and ML training have leaned strongly towards the weights in general not being considered a derivative work, but I have doubts that the law would see this as black and white; for example, even if the general consensus is that ML training to produce weights, in and of itself, does not create a derivative work, if you are able to show that a given set of weights is able to verbatim reproduce inputs (as a result of overfitting or memorization), I have my suspicions that it would not be shrugged off so easily. In true "color of my bits" fashion, I think that from a legal standpoint, the actual technical means by which something was accomplished doesn't matter if the process as a whole is effectively copyright infringement.

There do seem to be some ongoing cases regarding this such as Getty Images v. Stability AI and it will be interesting to see their result.

Re: OpenAI: Copy, Steal, Paste

#35

Can someone explain how adjusting weights based on viewing data is copyright infringement? I don't know if it's just not understood how these models work, or if it's just purposely misleading to try and cripple scary new tech. It seems like writers trying to complain it's copyright infringement to have other writers read their works for inspiration.

This argument, like everything else, is meaningless unless we talk about scale. It's like saying, well, I don't mind if someone sits by the road with pen and paper writing down license plate numbers.

But now, when the scenario is adjusted to be a network of ALPN sensors blanketing an entire metro area, am I just "purposely misleading to try and cripple scary new tech" when I object to such a system?

Re: OpenAI: Copy, Steal, Paste

#36

Technically it's not a copy. And nothing was stolen. It's a best fit curve among a series of datapoints. The datapoint is a copy, if the best fit curve never touches the datapoint then it's technically not a copy. The difficult part is the technicality here is legal and ethical from any standpoint. The high level ramifications are a bit unfair in the sense that yes the data is being used to create an AI that can repl…

Copyright laws do not support your argument.

There have been many cases in music where the offending song was forced to pay because it was "close enough" to the curve but not touching it.

Re: OpenAI: Copy, Steal, Paste

#37
post #23
post #2

Can't agree more. These AI systems have been built on the back of free labor for decades. We will look back on this and wonder how we essentially subsidized these behemoth corporations, then allowed them to extract subscription fees back from the very people it took the labor from. Don't get me wrong, there are genuine innovations in AI and ML. But, on the same token, you can't have ChatGPT without content.

> These AI systems have been built on the back of free labor for decades. Do you repay publishers for information you’ve summarized or new insights you’ve gained after consuming their content? How is it different when AI does the same thing?

I paid for the newspaper, book copy or museum visit from which I learnt and got inspired. And like me, million of people did over the years. And yes, I also pirated part of that content as well.

Re: OpenAI: Copy, Steal, Paste

#38

I don't think it should be ok to copy the literal works, the example of it emitting large chunks of articles from NYT is egregious, but it should be ok for it to learn and synthesize anything it can see. It should ideally, for producers and consumers of information, also be able to function as an index. Some of the pushback is the desire to collect fees or social capital for works, which is understandable, but at the…

It is easy to check for literal copies of the training material. I expect the providers to just add a filter doing that.

But I don’t think it resolves all the problems. Newspapers are losing their business model. Maybe it is just creative destruction and we’ll have something better in their place. Maybe, but journalism is quite important in our societies.

Re: OpenAI: Copy, Steal, Paste

#39
post #26
post #2

Can't agree more. These AI systems have been built on the back of free labor for decades. We will look back on this and wonder how we essentially subsidized these behemoth corporations, then allowed them to extract subscription fees back from the very people it took the labor from. Don't get me wrong, there are genuine innovations in AI and ML. But, on the same token, you can't have ChatGPT without content.

In a just world, royalties would be paid to those whose content AIs were trained on. In reality, the most we'll likely see is a few people who successfully sue for their small piece of the pie, while the big companies carry on like nothing happened - further accelerating the transfer of wealth right to the top.

In that just world, the royalties would be pennies, just as is the case with Spotify. Creators would then endlessly complain that their content isn't worth more than it is.

It's fine if we shift to that model. The content creators won't be well compensated in that world, because their individual pieces of content are not in fact particularly valuable. It's a delusion that they are. What they really want is an artificially inflated over-compensation for their share (and if everyone got that, it wouldn't work at all).

And exactly as with Spotify, even with the payouts, the big corporation (OpenAI) will be the valuable thing in the end.

Re: OpenAI: Copy, Steal, Paste

#40
post #26
post #2

Can't agree more. These AI systems have been built on the back of free labor for decades. We will look back on this and wonder how we essentially subsidized these behemoth corporations, then allowed them to extract subscription fees back from the very people it took the labor from. Don't get me wrong, there are genuine innovations in AI and ML. But, on the same token, you can't have ChatGPT without content.

In a just world, royalties would be paid to those whose content AIs were trained on. In reality, the most we'll likely see is a few people who successfully sue for their small piece of the pie, while the big companies carry on like nothing happened - further accelerating the transfer of wealth right to the top.

In a just world, publishers would be able to opt-out of AI training, and AI companies would have to respect this. Then, if they still want the material, they can approach publishers and buy the content. (Smaller publishers might make a "content pool" so they don't have to negotiate individually)

(This is how current search engines work, btw.. and that's why no one complains about them crawling content)

Post reply on HN