Live data from Hacker News

AI is just unauthorised plagiarism at a bigger scale

axelk.ee

631–640 of 783 posts

Re: AI is just unauthorised plagiarism at a bigger scale

#631

Earlier quoted context omitted.

It's not like that, because flowers are a physical object and moving them to one place deprives their original location of the flowers. When an LLM learns something from a webpage, the webpage is still there. Whatever 'theft' I perceive is entirely in my head; I was deprived of nothing by someone else making a copy of my thing.

I get that the intention here is to plagiarize and thus cause the parent to feel the harm of it and realize the error in their ways, but I don't think it works. Plagiarism's harm to the plagiaree (?) is that it robs them of credit and payment, but nobody is viewing your reply in isolation of the parent's attribution and parent wasn't expecting to make money off of an HN comment. The harm to the rest of society where…

> google has hastily indexed this

Google doesn't claim authorship over that which they index.

Plagiarism doesn't need to be harmful for it to be bad, and my intent wasn't to harm anyone anyway. My intent was that I could use the authors exact words to pretend to make a unique take that I claimed to have authored.

Re: AI is just unauthorised plagiarism at a bigger scale

#632
post #82
post #5

People were effectively copying websites (especially ecommerce tutorials) and beating the original authors at SEO decades before ChatGPT 2.

People also got blown up before atomic bombs, but it's hard to argue that they weren't worth treating more seriously than a stick of dynamite. Sometimes being able to do something at a massively larger scale is a meaningful difference.

But ChatGPT does nothing to scale copying somebody else's website. What are we talking about here, exactly? The article doesn't link to the original or clone, they don't mention rephrasing, and they specifically call out link text being the same. Even if you need the cloned site not to be identical, a thesaurus + scraper should scale far better than having an LLM do it!

Re: AI is just unauthorised plagiarism at a bigger scale

#633
post #519

Earlier quoted context omitted.

I assume he's saying Disney owns the 1992 film so the 1999 film is not theft, but he wants it to be because he doesn't like the 1999 film. Thus the quotes.

That's not a charitable reading of the comment, and furthermore, it's not even a reasonable assumption. Other comments clarify that the "theft" is in quotes because it's a figurative theft, not from Disney to themselves, but from Disney to the earlier, non-copyrighted folk tales it drew inspiration from. And the "theft" is that the Disney IP supplanted (via ubiquity) the public domain versions to the point lots of pe…

I don't think it's "uncharitable"? Seems perfectly reasonable to not like a remake.

He says:

> ... this corporate remake is a worse creative "theft" than ...

Context is that "this" is the 1999 film.

A sibling comment makes a separate point that even the 1992 film is not original content but nowhere in falcor84's comment does he refer to the franchise as a whole being "theft".

Regardless, it's clear from the post that the context is the 1999 film being `creative "theft"` which I inferred meant they changed the story in ways he didn't like but... he can weigh in if he feels like it.

Re: AI is just unauthorised plagiarism at a bigger scale

#634
post #286
post #252

Earlier quoted context omitted.

Then go ahead and abolish copyright for everyone. Instead we're stuck in an even worse system where the hypercorporations gleefully plagiarize everyone else while sending SWAT teams to kill anyone who pirates a movie.

Obviously there's an ideal middle ground, but what LLMs do is allow free transfer of knowledge while still (mostly) preserving the protections that copyright should be protecting. For example, I can have an LLM give me the entire plot of a book (which is fine), but it won't spit out an exact copy of the book.

I can get an LLM to spit out an exact copy of my python-docx library.

Re: AI is just unauthorised plagiarism at a bigger scale

#635
post #169

This is really not so clear cut as "fair use" might cover 99% of all data scrapping; you are not reproducing the originals just use them to estimate probabilistic distribution of tokens in pre-training. You are never going to get the exact book word-for-word using LLMs.

This confuses input and output. A copy made for the purposes of training is still a copy. Even if you throw the text away after training, you've still made a copy.

In Bartz v. Anthropic the judge ruled that Anthropic making a digital copy of a printed book and then discarding the physical book was not infringing when used to a train a model.

Re: AI is just unauthorised plagiarism at a bigger scale

#636
post #464

The linked article shows that LLMs can be used to plagiarize content through rewriting. Then he gets SEO'd out of it. But it doesn't demonstrate that AI is just plagiarism.

It doesn't even show that LLMs can plagiarize through rewriting, is just asserts that. The audience is expected to already believe it. And, fair enough, it's true. But I can't make any sense of 700+ upvotes for what I'd say amounts to a 200-word disjointed HN rant comment if it even met site guidelines.

Re: AI is just unauthorised plagiarism at a bigger scale

#637
post #536

Earlier quoted context omitted.

The conversion from viewer to donator is around 1%. This is true from wikipedia, to twitch, to podcasts. The number of people who will not ever load your ads is around 30%. I can tell you that creators talk about this a lot in private, but will not publicly because the internet has a mass delusion on how creation and compensation works. It's like trying to convince christians that jesus obviously didn't come back fro…

The problem is that the ad vendors couldn't keep it in their pants. The ads you're talking about are a common vector for delivering malware onto people's PCs, and absolutely destroy the usability of sites. Between tracking cookies, popups, full screen banners, autoplaying video, flashing ads, and their unbelievably high weight in bandwidth - the internet is fairly unusable if you don't block any ads Bear in mind that…

>The problem is that the ad vendors couldn't keep it in their pants.

You might not know, many people don't, that ad vendors came to the table little over a decade ago to make a truce with Ad Block Plus. ABP and advendors both saw that an "ad supported internet" was unsupported with no ads. So ABP was looking to set terms for what would be deemed as acceptable ads. Creators/service providers get incentive, users get manageable ads.

It didn't matter though because users rioted and uBlock (then uBlock Origin) became king. No compromises there. I mean, what fucking idiot would take some ads when they could take no ads, right?

Even less known is that Google trailed a program where you could pay them directly and they would remove ads from your browsing. This program was about as popular as shit on stick, because again, what fucking idiot would pay for no ads when they simply block all ads for free, right?

There have also been attempts like Brave, where crypto could be used as a micropayment in lieu of ads. But that has also gone nowhere, even if it does have a few snags around centralization.

What I have never seen though, and have zero examples of, is internet users trying to reconcile the situation. It's just a relentless entitlement to free everything, with a small fraction sometimes subscribing, and an even smaller fraction sometimes donating. The users are unquestionably the biggest assholes in this situation. They won't even acknowledge they have a problem.

Re: AI is just unauthorised plagiarism at a bigger scale

#638

Earlier quoted context omitted.

I get that the intention here is to plagiarize and thus cause the parent to feel the harm of it and realize the error in their ways, but I don't think it works. Plagiarism's harm to the plagiaree (?) is that it robs them of credit and payment, but nobody is viewing your reply in isolation of the parent's attribution and parent wasn't expecting to make money off of an HN comment. The harm to the rest of society where…

> google has hastily indexed this Google doesn't claim authorship over that which they index. Plagiarism doesn't need to be harmful for it to be bad, and my intent wasn't to harm anyone anyway. My intent was that I could use the authors exact words to pretend to make a unique take that I claimed to have authored.

I don't understand. In what way is plagiarism bad if it doesn't harm? If it were harmless to pretend you authored a unique take, how is the parent expected to react to you not harming them such that they realize it's bad?

Re: AI is just unauthorised plagiarism at a bigger scale

#639

There’s a fallacy that gets used a whole lot to justify things like this (not just with LLMs), and I see it in many of the comments here: If it’s OK (or at least negligible on a small scale), then it must be OK on a large scale. It usually goes something like: If I can make money by learning something from a web page, why does a computer making money by learning everything from everyone upset people so? It’s the same…

> It’s like if I go to Golden Gate Park and pick one flower, I shouldn’t do that, but no one cares. But if I build a machine to automatically cut every flower in the park because I want to sell them, that’s different. The problem here is, in your example the small scale example, and the large scale example are both unacceptable behavior. Learning from others at a small scale is not only socially acceptable, but is th…

> Learning from others at a small scale is not only socially acceptable, but is the foundation of how advancement works.

Exactly, if anything, the logic (a bit bad -> really bad) shows that one person learning from one thing is far inferior to one person learning from every thing (a bit good -> really good).

Post reply on HN