Live data from Hacker News

AI is just unauthorised plagiarism at a bigger scale

axelk.ee

671–680 of 783 posts

Re: AI is just unauthorised plagiarism at a bigger scale

#671

Earlier quoted context omitted.

This is a great point. I think for coding, the wording of the MIT open source license makes it clear that copying and distributing the software is authorised on a small scale and it's very clear that the act of copying must involve a person. It provides distribution and modification rights to "any person obtaining a copy of the software" and explicitly requires attribution for any significant parts. Mass-ingesting th…

> I think for coding, the wording of the MIT open source license makes it clear that copying and distributing the software is authorised on a small scale and it's very clear that the act of copying must involve a person. I agree with “must involve a person. https://opensource.org/license/mit starts with (emphasis added) “Permission is hereby granted, free of charge, to any PERSON obtaining a copy of this software and…

The CI pipeline is different because for a module to end up as a dependency in the CI pipeline, it had to be explicitly selected by a person first to be included in the package file or manifest. There was intentionality and awareness that the software was included.

A person already pre-consented to the licenses of all the software which the pipeline downloaded. Big companies go through those dependency lists carefully already and remove those which do not meet their policies. This is a very intentional process.

Re: AI is just unauthorised plagiarism at a bigger scale

#672

Earlier quoted context omitted.

But in this case, a human has awareness of what software they are copying or modifying and that's how the original software author receives credit. The contract requires some degree of human awareness to be valid. This is the critical difference.

Sorry that's nonsense. There's human awareness when ingesting MIT code into an LLM too. In both cases it's a human that says $ excute-global-replace or $ ingest-into-llm Both operations require some degree of human awareness. What you appear to be saying is, a human can only use a limited algorithm to access this source code, not a sophisticated one. And where do you draw that line? Who should get to say what is too…

If your LLM were to hack into Microsoft and steal the source code from an important project and inject it into your project without you being aware of it; wouldn't that make you liable if you then published it?

Unfortunately there is no way to agree to a license of a software you're using if you didn't read the license or if you're not even aware that you're using the licence. This is what's happening at the training stage.

If you say that awareness doesn't matter then it means you cannot stop AI from stealing any IP open source or not.

I think the main issue with LLMs is that there is no mechanism to stop them from stealing. Thus they are guaranteed to infringe on copyright to some extent.

Also, beyond copying and copyright, there is another problem that LLMs are also infecting the logic and expertise built into the project. This is a completely novel mechanism and needs to be treated as separate under the law. Else it would be the end of all IP.

Re: AI is just unauthorised plagiarism at a bigger scale

#673

There’s a fallacy that gets used a whole lot to justify things like this (not just with LLMs), and I see it in many of the comments here: If it’s OK (or at least negligible on a small scale), then it must be OK on a large scale. It usually goes something like: If I can make money by learning something from a web page, why does a computer making money by learning everything from everyone upset people so? It’s the same…

But also a person is a person, not a commercial product, and usually we learn from sources within their licensing agreements.

Re: AI is just unauthorised plagiarism at a bigger scale

#674

There’s a fallacy that gets used a whole lot to justify things like this (not just with LLMs), and I see it in many of the comments here: If it’s OK (or at least negligible on a small scale), then it must be OK on a large scale. It usually goes something like: If I can make money by learning something from a web page, why does a computer making money by learning everything from everyone upset people so? It’s the same…

> There’s a fallacy that gets used a whole lot to justify things like this ...

FWIW, this is the Fallacy of Composition

https://en.wikipedia.org/wiki/Fallacy_of_composition

Re: AI is just unauthorised plagiarism at a bigger scale

#675

Earlier quoted context omitted.

....Yet another vector through which "security experts" has caused a waterbed problem. Let's secure the Internet, oh no! We made a centralized list of operating domains for hostile actors to guide attacks with!

Sure, let's hide everything behind obscure schemes which will definitely serve the spirit of openness of the web.

The point is that you can't escape side-channel applications of security metadata being weaponized the more you try to force ubiquity of "security" everywhere. As long as there are motivated, profit seeking attackers, you have to take into account the toxic nature of metadata. This is another example of "A System Is What It Does" proving the pointlessness of "POSIWID". Intent doesn't matter. Certificate transparency was intended to clue us into bad cert issuing, but it is also a list of potential targets where AI crawlers can be directed to scrape new data. Intent doesn't change what it is. Cert transparency is certainly transparency + a "training data might end end up here" list.

Re: AI is just unauthorised plagiarism at a bigger scale

#676

There’s a fallacy that gets used a whole lot to justify things like this (not just with LLMs), and I see it in many of the comments here: If it’s OK (or at least negligible on a small scale), then it must be OK on a large scale. It usually goes something like: If I can make money by learning something from a web page, why does a computer making money by learning everything from everyone upset people so? It’s the same…

The era before internet, the gaps among information and knowledges could make money and power.

The era after internet and before LLM, the information and knowledge gaps have been largely leveled theoretically, but the recognition wall stops most of us to understand and make use of them.

The era after LLM, the wall is being destroyed and people should think about how to use these information and knowledge differently to make money and power.

Re: AI is just unauthorised plagiarism at a bigger scale

#677
post #268

Earlier quoted context omitted.

I don't disagree with your premise, but I'd argue that saying "there are no original ideas" in the context of a discussion of plagiarism is needlessly reductive. Even though I think I mostly agree with the author here, I think there are legitimate counterarguments that can be made; equating all of the ways someone can cite or build upon an idea with copying something word-for-word and claiming it's your own is not on…

No offense, but you sound like someone who has never built a language model. Anyone who has actually built one understands that there is no copying going on. Just predicting words (tokens actually). The problem is that people's words are MUCH more predictable then they would like to believe. And that truth upsets them. In addition to having created models, I also write books and articles. Probably more than most peop…

> No offense, but you sound like someone who has never built a language model. Anyone who has actually built one understands that there is no copying going on. Just predicting words (tokens actually).

> The problem is that people's words are MUCH more predictable then they would like to believe. And that truth upsets them.

I'm not offended. I do think it's a little weird that you seem to think "training on a bunch of stuff that includes a set of words" and then "predicting" those words exactly is somehow okay because theoretically it might be extrapolating the exact same words from combining other ones. I'd argue that if a model trains on data, and then reproduces exactly a large subset of that data, the bar should be pretty high to prove that it's not copying, and "you don't understand because you didn't implement this" is not a good basis for law.

> In addition to having created models, I also write books and articles. Probably more than most people commenting here. I have a firm grip on what actual copyright law is and the pros and the cons of it.

I'm not convinced you have a firm grip on the idea that no matter how smart you may be, "just trust me bro" is a pretty terrible strategy if you're actually intending to convince anyone of anything. If that's not what your goal is here, it's not clear why it's worth your time to respond to other people's comments when you clearly have so many other productive ways to spend your time.

Re: AI is just unauthorised plagiarism at a bigger scale

#678
post #91

Earlier quoted context omitted.

When did the last original thought happen then? Clearly thoughts must have been original at some point, or there wouldn't be any at all

You seem to discount the possibility that ideas are emergent and as they emerge, multiple people at once become aware of them. I am asserting it is Charles Fort's "steam engine time". Far from a crank position. It is one that bears serious consideration.

No, I'm saying that simultaneous discovery and plagiarism are not philosophically incompatible, and treating them as equivalent is hard to take seriously.

Re: AI is just unauthorised plagiarism at a bigger scale

#679
post #82

Earlier quoted context omitted.

People also got blown up before atomic bombs, but it's hard to argue that they weren't worth treating more seriously than a stick of dynamite. Sometimes being able to do something at a massively larger scale is a meaningful difference.

But ChatGPT does nothing to scale copying somebody else's website. What are we talking about here, exactly? The article doesn't link to the original or clone, they don't mention rephrasing, and they specifically call out link text being the same. Even if you need the cloned site not to be identical, a thesaurus + scraper should scale far better than having an LLM do it!

> But ChatGPT does nothing to scale copying somebody else's website. What are we talking about here, exactly?

"Make me a website that has the same content as that other one so I can get views instead" is not something you could could do generically and quickly with a free service a few years ago, but it is today. I'd argue that it's not beneficial to people who create original content or society at large that this is the case. There are plenty of other uses of LLMs, some of which are genuinely beneficial, some which are mixed, and some which are also a net negative. It seems pretty reasonable to me that issues like this are worth discussing, because as all of the comments on this article here show, people clearly are not on the same page about it.

Re: AI is just unauthorised plagiarism at a bigger scale

#680

Earlier quoted context omitted.

>Why look at a website when it's all in AI? well, at least in the case of google, I'm pretty sure that's the point. Or at least, they are doing things that would seem to be moving towards being an oracle with all the answers and not the signpost that points you in the right direction. The destination rather than the gateway.

remember AMP?

Holy cow, I never even thought about that in relation to AMP. It's not a new thing, then.
Post reply on HN