Live data from Hacker News

AI is just unauthorised plagiarism at a bigger scale

axelk.ee

721–730 of 783 posts

Re: AI is just unauthorised plagiarism at a bigger scale

#721
post #169

This is really not so clear cut as "fair use" might cover 99% of all data scrapping; you are not reproducing the originals just use them to estimate probabilistic distribution of tokens in pre-training. You are never going to get the exact book word-for-word using LLMs.

Fair use was built around human limitations. The mass scraping campaigns done by the AI giants were clearly an overreach in spirit, if not letter. Most people's intuition is that these massive operations that are valued in the trillions can't have been drawn from some untapped common resource, and they're correct. Someone, somewhere is not being properly compensated. I have no problem with taxing AI companies so that…

Fair use is the balance between creators and those that in someway use the content. Somehow it has become excuse not to compensate the creators in anyway. To me AI training part really looks something that should be treated separate and thus give the creators compensation when their works are used.

Now how much and should it be based on revenue from output is open discussion. And it might also be that there is no fair model to pay them. Which means that well too bad for LLMs...

Re: AI is just unauthorised plagiarism at a bigger scale

#722
post #171

if theres just one good thing coming out of ai its breaking copyright law forever. no one should be able to "own" ideas. royalties for commercial use is another thing and i support it but what we know as (non commercial) piracy and unlicensed fan art should be 100% legal

I wonder how many of the books I love would still have been written in a world where somebody could scoop them all up and post them on the internet for free (and run ads).

I feel not that many. Or at least many successful authors would struggle lot more if after launch of new book next week anyone could be selling poorly made cheaper copy in stores.

And most likely ones doing that would be your biggest companies say Amazon.

Re: AI is just unauthorised plagiarism at a bigger scale

#723

Earlier quoted context omitted.

This is how the drug industry already works. I don’t think there’s any evidence “AI” (LLM) is capable of producing valid drug modifications.

In current status AI models cannot do that. But, if they do then it will break Medical Patent model.

The value in Medical patents is not the idea. It is the process of proving efficacy and safety. Which are the expensive parts. And I doubt we will trust AI with those any time soon. We grand Medical Patents because proving things is expensive and that process needs to be encouraged.

Re: AI is just unauthorised plagiarism at a bigger scale

#724

Earlier quoted context omitted.

I could say the same of your position, honestly. Stupid, naive - or maybe just plain ignorant. If humans didn't want information to be free, there wouldn't be so much free information. Or did you not notice?

You are confusing "slop" with "information", there is so much slop because it costs nearly 0 to be produced, but there's far less "information" than you are thinking.

Slop is just information that a human can't be bothered to process, for a near-infinity multitude of reasons.

Re: AI is just unauthorised plagiarism at a bigger scale

#725
post #591

There’s a fallacy that gets used a whole lot to justify things like this (not just with LLMs), and I see it in many of the comments here: If it’s OK (or at least negligible on a small scale), then it must be OK on a large scale. It usually goes something like: If I can make money by learning something from a web page, why does a computer making money by learning everything from everyone upset people so? It’s the same…

> But quantitative changes in an activity produce qualitative changes. Interesting take. I think a corollary is that the qualitative changes are in the economics of things. And more than the scale, it is the value of those economic effects that determines how "accepted" that activity becomes. Take Uber as an example; it basically enabled mass avoidance of taxi regulations, and naturally existing taxi drivers and lawm…

> gradually and inexorably society and laws adjusted to it.

But in many places, the ways that society and laws adjusted to it were to make extra clear in their local ordinances that Uber was required to operate as an actual taxi service, or get out.

It's very disingenuous to imply that the public broadly decided Uber was Right, Actually, when both in its case and in that of many of the other gig economy companies, what really happened is that gradually and inexorably, they had to adjust to society and laws.

Re: AI is just unauthorised plagiarism at a bigger scale

#726

Earlier quoted context omitted.

Sorry that's nonsense. There's human awareness when ingesting MIT code into an LLM too. In both cases it's a human that says $ excute-global-replace or $ ingest-into-llm Both operations require some degree of human awareness. What you appear to be saying is, a human can only use a limited algorithm to access this source code, not a sophisticated one. And where do you draw that line? Who should get to say what is too…

If your LLM were to hack into Microsoft and steal the source code from an important project and inject it into your project without you being aware of it; wouldn't that make you liable if you then published it? Unfortunately there is no way to agree to a license of a software you're using if you didn't read the license or if you're not even aware that you're using the licence. This is what's happening at the training…

> I think the main issue with LLMs is that there is no mechanism to stop them from stealing.

Well, sure there is—for the people running them.

If you're building training data for an LLM, you only use data that a) is firmly in the public domain, or b) you have a clear and documented legal right to use.

Re: AI is just unauthorised plagiarism at a bigger scale

#727

There’s a fallacy that gets used a whole lot to justify things like this (not just with LLMs), and I see it in many of the comments here: If it’s OK (or at least negligible on a small scale), then it must be OK on a large scale. It usually goes something like: If I can make money by learning something from a web page, why does a computer making money by learning everything from everyone upset people so? It’s the same…

> why does a computer making money by learning everything from everyone upset people so? It’s the same thing! The majority of the population, sitting outside the VC bubble, views AI unfavorably. That's not my hot take, that's a fact from the NYT survey published today. It's going to be hilarious when VCs, having expropriated the IP of the entire internet, build The Layoff Machine That Does Everything Without Workers,…

The problem is, there's an intermediate step required there: the voters will need to get rid of the Republican Party lock, stock, and barrel if they ever want to make a genuinely-socialist move like that. And that's going to be made much, much more difficult by all the measures Trump and his cronies are putting in place to disenfranchise everyone who refuses to bow to him.

Re: AI is just unauthorised plagiarism at a bigger scale

#728

Earlier quoted context omitted.

> why does a computer making money by learning everything from everyone upset people so? It’s the same thing! The majority of the population, sitting outside the VC bubble, views AI unfavorably. That's not my hot take, that's a fact from the NYT survey published today. It's going to be hilarious when VCs, having expropriated the IP of the entire internet, build The Layoff Machine That Does Everything Without Workers,…

>The majority of the population, sitting outside the VC bubble, views AI unfavorably. Sure, where AI means threatens my job or my skills, people view it unfavourably. But then they use it. They're all using it. People's rhetoric seldom matches their actions. >enthusiastically expropriate that, and we end up with Fully Automated Luxury Communism Maybe in other countries, initially, but the US is very firmly a plutocra…

> They're all using it.

[Citation needed]

I know many more people who do not use AI than who use it, and many more who refuse to use AI than people who are enthusiastic about it.

Given your username, you are almost certainly in a bubble—an echo chamber—that makes it seem to you as though "everyone is using it." I recommend getting outside that bubble and talking to non-technical people outside your usual circles, especially people in the arts and humanities.

Re: AI is just unauthorised plagiarism at a bigger scale

#729

Earlier quoted context omitted.

Sure, let's hide everything behind obscure schemes which will definitely serve the spirit of openness of the web.

The point is that you can't escape side-channel applications of security metadata being weaponized the more you try to force ubiquity of "security" everywhere. As long as there are motivated, profit seeking attackers, you have to take into account the toxic nature of metadata. This is another example of "A System Is What It Does" proving the pointlessness of "POSIWID". Intent doesn't matter. Certificate transparency…

[deleted]

Re: AI is just unauthorised plagiarism at a bigger scale

#730

Earlier quoted context omitted.

> I think for coding, the wording of the MIT open source license makes it clear that copying and distributing the software is authorised on a small scale and it's very clear that the act of copying must involve a person. I agree with “must involve a person. https://opensource.org/license/mit starts with (emphasis added) “Permission is hereby granted, free of charge, to any PERSON obtaining a copy of this software and…

The CI pipeline is different because for a module to end up as a dependency in the CI pipeline, it had to be explicitly selected by a person first to be included in the package file or manifest. There was intentionality and awareness that the software was included. A person already pre-consented to the licenses of all the software which the pipeline downloaded. Big companies go through those dependency lists carefull…

> for a module to end up as a dependency in the CI pipeline, it had to be explicitly selected by a person first

I disagree. I think it’s entirely within the license to have your pipeline automatically pull in the latest version of a library, even if the new one happens to pull in a new MIT-licensed library (whether that’s a good idea and whether CI pipelines should, somehow, verify that code pulled in has an acceptable license are different discussions)

I also think it’s complete within the MIT license to tell a LLM that it can search for MIT-licensed libraries and use them without asking you.

Post reply on HN