Live data from Hacker News

OpenAI: Copy, Steal, Paste

computerworld.com

71–79 of 79 posts

Re: OpenAI: Copy, Steal, Paste

#71
post #46

Earlier quoted context omitted.

TensorFlow is also perfectly legal to develop and distribute, and no one contests this. People object to specific artifact, "model weights", which were produced using copyrighted works at the input, and can be used to reproduce those same copyrighted works back. In bittorrent analogy, people want to shut down specific pirate trackers and the pirate bay website.

One thing that I think people forget about is that the prompt used when "reproduc[ing] those same copyrighted works" is also a part of why it spits out similar things. It's not just the model doing it. A traditional artist can be prompted to recreate a copyrighted work in much the same way with the right prompts.

I don't think most people are misinterpreting things. The truth is that models which are not terribly overfit literally don't output verbatim inputs often, in fact, for Stable Diffusion it's apparently nearly infinitesimally small odds, and this is good because that implies that the weights are in fact, not literally encoding some crazy kind of compressed copies of the images in question.

On the other hand, if you prompt a code generating model with some comment and a function declaration that it knows exists and it spits out 100+ lines of nearly verbatim code, that's a completely different story entirely. If I prompt a human with that sort of thing, they will almost certainly write different code even if they've seen the original source code in question. This is in part because the way humans write code is different from the way LLMs write code; humans tend to iterate somewhat non-linearly, and I think if you ask the same person to write the same thing on different days, they would probably come up with different results. It would be quite rare for a human to just see a familiar segment of code and then begin dumping near-verbatim copies of existing codebases.

AI models that readily and easily bias themselves toward outputting their inputs do exist. It is not clear how many models actually do this, but this is definitely a huge part of the concern when people talk about copyright and model weights.

It's a bit clouded by people who are just generally hoping that today's AI model weights are illegal for social reasons, but that's not the position I am trying to present. (I'm not really sure what we should do regarding societal impact.)

Re: OpenAI: Copy, Steal, Paste

#72

Earlier quoted context omitted.

Copyright laws do not support your argument. There have been many cases in music where the offending song was forced to pay because it was "close enough" to the curve but not touching it.

I think it's an apt analogy, though I disagree about the implication. If I use ChatGPT to create a work, and that work is "close enough" to an existing copyrighted work, then it seems like I am guilty of copyright violation, not ChatGPT.

It's not an analogy. This is actually what is done with ML. It is literally a best fit curve problem.

Or maybe it is actually an analogy, but then if this was the case the entire field of ML is capable of only understanding the intricacies of ML through the analogy of curve fitting and what's actually going on underneath the analogy remains elusive.

Re: OpenAI: Copy, Steal, Paste

#73
post #23
post #2

Can't agree more. These AI systems have been built on the back of free labor for decades. We will look back on this and wonder how we essentially subsidized these behemoth corporations, then allowed them to extract subscription fees back from the very people it took the labor from. Don't get me wrong, there are genuine innovations in AI and ML. But, on the same token, you can't have ChatGPT without content.

> These AI systems have been built on the back of free labor for decades. Do you repay publishers for information you’ve summarized or new insights you’ve gained after consuming their content? How is it different when AI does the same thing?

Are you implying normal humans get access to the data used in training GPts for free? Where can I read all those books for free, from an offline copy and without ads?

Re: OpenAI: Copy, Steal, Paste

#74

Can someone explain how adjusting weights based on viewing data is copyright infringement? I don't know if it's just not understood how these models work, or if it's just purposely misleading to try and cripple scary new tech. It seems like writers trying to complain it's copyright infringement to have other writers read their works for inspiration.

Adjusting weights based on viewing data also describes a JPEG file, if you squint. "I didn't copy a picture, I merely adjusted cosine transform weights after viewing it."

Since you asked for someone to explain: IMO you're focusing on a mechanistic explanation of how ML happens, but if you take only a step back, the resulting models are really similar to a compressed blob of the training data. At times they regurgitate the training data almost exactly. This is what most people are seeing and (IMO, rightly) calling out as blatant infringement.

Slightly different topic: I think calling it "learning" is brilliant PR, but if you survey AI experts, they'll mostly agree it's not learning in the same way a human brain learns. You and I get more legal leeway than a computer, and it'll be interesting how that works out if/when AGI becomes a thing, but we're nowhere near there yet.

Re: OpenAI: Copy, Steal, Paste

#75
post #12

Earlier quoted context omitted.

>These AI systems have been built on the back of free labor for decades. In practical terms, is that different from Google becoming a trillion-dollar company from its search offering?

Indexing content provides a useful service that content owners benefit from, too, so for a long time, there was definitely a mutual understanding between Google and site owners. I think the relationship has soured, somewhat coincidentally, in large part due to the way that Google has started to push against actually linking to sites, like by trying to provide inline answers and using AMP caches to avoid actually drop…

>Indexing content provides a useful service that content owners benefit from

I'm not sure about that. The Google Index put a lot of existing content creators out of business, like local newspapers. They certainly did not see this great benefit of a Google index. Even today, given the insanely low ad rates, because it's impossible to generate income from public content (hence the rise of paywalled subscription services) no content creator is making any money - and yet, Google is still worth a trillion dollars. And yes, there's also the fact that in the last decade or so, Google has started side-stepping the content providers by simply providing 'inline answers' and 'AMP caches'

So no, I don't see much difference between Generative AI and your 'traditional' aggregation services.

Re: OpenAI: Copy, Steal, Paste

#76
post #12

Earlier quoted context omitted.

Indexing content provides a useful service that content owners benefit from, too, so for a long time, there was definitely a mutual understanding between Google and site owners. I think the relationship has soured, somewhat coincidentally, in large part due to the way that Google has started to push against actually linking to sites, like by trying to provide inline answers and using AMP caches to avoid actually drop…

>Indexing content provides a useful service that content owners benefit from I'm not sure about that. The Google Index put a lot of existing content creators out of business, like local newspapers. They certainly did not see this great benefit of a Google index. Even today, given the insanely low ad rates, because it's impossible to generate income from public content (hence the rise of paywalled subscription service…

[deleted]

Re: OpenAI: Copy, Steal, Paste

#77
post #12

Earlier quoted context omitted.

Indexing content provides a useful service that content owners benefit from, too, so for a long time, there was definitely a mutual understanding between Google and site owners. I think the relationship has soured, somewhat coincidentally, in large part due to the way that Google has started to push against actually linking to sites, like by trying to provide inline answers and using AMP caches to avoid actually drop…

>Indexing content provides a useful service that content owners benefit from I'm not sure about that. The Google Index put a lot of existing content creators out of business, like local newspapers. They certainly did not see this great benefit of a Google index. Even today, given the insanely low ad rates, because it's impossible to generate income from public content (hence the rise of paywalled subscription service…

> The Google Index put a lot of existing content creators out of business, like local newspapers. They certainly did not see this great benefit of a Google index.

I dunno why you ignored the entire part of my comment wherein I describe the gradual fallout with Google, but frankly it's difficult to even begin with how infuriatingly wrong this line of thinking is.

So basically, you're saying that local newspapers, for example, did not see some great benefit from the Google search index. In fact, they're being killed by it! Shivers.

But wait a minute. How? How is Google killing the websites that its indexing? Is it because they're causing their hosting bills to skyrocket through negligent web crawling? Doesn't seem to be. Is it because they're charging websites through the nose to be listed? No, being listed in Google is free... Yeah, the reason why Google is "killing" local newspapers is simply local newspapers don't make as much money from Google traffic as they used to.

Forget about how Google is benefiting from local newspapers without giving back to them as much. That is happening, but there is a crucial and specific detail in your argument that needs to be scrutinized. If websites simply saw no benefit from being indexed by search engines, uhm, They Simply would choose to not be indexed in search engines. You can choose to do that. But they won't, because the problem isn't actually that Google is killing them. That's just the soppy rhetoric used to sell asinine ideas like the Link Tax wherein you will be billed for providing a hyperlink to a website.

The point I was going on about in my original comment was in fact that search indices like Google were in fact, always a give-and-take operation. Google benefits from having content in the index, websites benefit from the traffic they get from Google. The difference today is that Google does work harder to keep you off of the actual target website by adding features like info boxes containing snippets and again, AMP caches. However, note that these features don't really increase the "take": Google is taking exactly the same thing it always did, it's indexing your web sites and presenting its information in a transformed format on theirs. What it's giving, however, in terms of traffic to websites, has decreased over time.

Meanwhile, generative AI is a take operation. Certainly both search engines and generative AI models provide utility to the users of their respective services, but in case of generative AI models, it's only in service of the owners of the model, who give essentially nothing back, not credit, not a link, and certainly not revenue sharing of any kind.

This isn't in defense of Google's practices by any means: Like I said, they upset the balance that was in place from unspoken agreements. However, it definitely is in defense of the truth, and the truth is a phrase like "Google kills local newspapers" is bullshit, and if it wasn't, they'd all be running away from the Google index as fast as possible. Instead, many of them closely follow the latest SEO tricks, still.

And I hate this narrative anyways, because if organic Google traffic and Google Adsense was the only thing funding local newspapers, they were never sustainable by any measure in the first place.

Re: OpenAI: Copy, Steal, Paste

#78
post #9

It seems obvious to me that a system in which intellectual property can be ingested into an artificial intelligence system and then provide no value to the original creator is unsustainable. The value of content that is generated through real-world effort like research and investigation will plummet towards zero. And then what? You can't just have a society built off of LLMs feeding each other. It also feels obvious…

The moment the rules are finalized, flocks of bright unprincipled people will rush to gamify them to turn a profit.

Is the implication that we should therefore not try to come up with good rules?

Re: OpenAI: Copy, Steal, Paste

#79
post #78

Earlier quoted context omitted.

The moment the rules are finalized, flocks of bright unprincipled people will rush to gamify them to turn a profit.

Is the implication that we should therefore not try to come up with good rules?

No, we should calibrate expectations accordingly and not assume that even the signatories to the rules won’t rush to exploit the loopholes the nanosecond they put their signatures there.

There should be a way to codify what “the spirit” of the rules is and that seeking a way to subvert it is against the rules too.

Post reply on HN