Live data from Hacker News

AI is just unauthorised plagiarism at a bigger scale

axelk.ee

511–520 of 783 posts

Re: AI is just unauthorised plagiarism at a bigger scale

#511
post #503

Earlier quoted context omitted.

We ran into a lot of stuff like this in the early days of the web. For example, there was a lot of information that was "public" in that anyone could go to the city courthouse and ask to see the documents. But it changed in nature when you could suddenly look up anyone in the country by typing their name in your browser.

we used to ship mass lists of addresses and phone numbers to people in each town and it was fine/appreciated.

You could also easily opt out with the single entity that shipped that information.

Re: AI is just unauthorised plagiarism at a bigger scale

#512

There’s a fallacy that gets used a whole lot to justify things like this (not just with LLMs), and I see it in many of the comments here: If it’s OK (or at least negligible on a small scale), then it must be OK on a large scale. It usually goes something like: If I can make money by learning something from a web page, why does a computer making money by learning everything from everyone upset people so? It’s the same…

If one person is murdered, that's bad. If a million people are murdered, that's war. If one word is stolen by AI, that's bad. If a million words are stolen by AI, that's business.

>If one word is stolen by AI, that's bad. If a million words are stolen by AI, that's business.

Where are all the instances of "one word" being "stolen by AI", and people getting mad over it?

Re: AI is just unauthorised plagiarism at a bigger scale

#514
post #429

Earlier quoted context omitted.

> I'm afraid, the essence is that is not. Re-sequencing content is not the same as synthesis Drawing different sources of information together into a single understanding is quite literally the definition of "synthesis" in this context. If that process is what you're referring to as "re-sequencing content", then it does fit the definition of "synthesis" in this discussion. If you're using the phrase "re-sequencing co…

If we're talking about concepts and communication, in text, I don't know what meaning of synthesis to apply (as long as there is meaning), other than the meaning this has had for centuries. I think, aggregation, extraction and emulgating is something else. The very purpose of text is to transfer meaning, concepts, observations and complex thoughts to human readers for them to process. And we have built a complex fram…

> I think, aggregation, extraction and emulgating is something else.

Aggregating information, extracting underlying concepts, and combining those concepts into a unified expression is indeed the vernacular meaning of "synthesis" applicable to this discussion.

"Emulgating" is not a conventional English word. Is it a misspelling of "emulating"? I ask because using the term "emulating" here would again represent an instance of question begging, i.e. implicitly asserting the position that what's being discussed is merely the paraphrasing of singularly sourced information, and not the unification of concepts expressed in multiple sources, which I again believe is the very thing we are debating.

> And we have built a complex framework around this and for this. The fact that many feel that this framework is violated should hint at there being a problem, a conceptual discrepancy.

I don't think there necessarily is a problem or conceptual discrepancy here, any more than there has been for all of the centuries that people have been debating epistemology. The problem here is the same as for humans, and reduces to a rationalism vs. empiricism debate. AI tools are pure rationalists, and are solely capable of reasoning. However, many people behave this way as well, and exhibit a rationalist epistomology, even having emotional entanglements with their axioms to the point that they'll bend over backwards to reject evidence that falsifies empirical conclusions drawn from those axioms.

My biggest fear from AI is not that it isn't capable of inductive reasoning -- that's all it's capable of, as I see it -- but rather that the fact that its reasoning has no empirical anchor will lead people who are mired in rationalist epistemology to accept its conclusions uncritically.

In other words, the danger doesn't come from the fact that AI has no semantic awareness, but that people using it aren't seeking semantic validation in the first place, which is a problem already pervasive in our society.

Re: AI is just unauthorised plagiarism at a bigger scale

#515

Earlier quoted context omitted.

https://en.wikipedia.org/wiki/Aladdin The argument, as I understand it is that the "theft" is in quotes because it's not literally copyright infringement, but fair use of an old public-domain folk tale that ends up consuming the latter. Today, when kids know "Aladdin" they know the copyrighted/trademarked Disney character, not the traditional folk tale- that's the "theft" that happened.

Would most kids around the world even know Aladdin if it wasn't for the Disney copyrighted movie?

Very likely yes. I was very familiar with this story, and other "Arabian" tales, well before Disney made the original animated version.

We also had Grimm's fairy tales, which I loved reading, and nowadays am reading to my daughter, to her delight. Yes, with beheadings and child-eating monsters and witches.

Re: AI is just unauthorised plagiarism at a bigger scale

#516

Did I miss where OpenAI plagerized the disproof of the planar unit distance problem from?

It would be one thing for someone to say "AI is enabling plagiarism at a bigger scale", but to say it's "just plagiarism", surely one needs to explain who exactly the unit distance breakthrough was plagiarized from.

Re: AI is just unauthorised plagiarism at a bigger scale

#517

There’s a fallacy that gets used a whole lot to justify things like this (not just with LLMs), and I see it in many of the comments here: If it’s OK (or at least negligible on a small scale), then it must be OK on a large scale. It usually goes something like: If I can make money by learning something from a web page, why does a computer making money by learning everything from everyone upset people so? It’s the same…

Yes absolutely, when automation increases the rate of something many orders of magnitude that often is a qualitative difference.

It's weird to me how often on HN of all places I see arguments that can be refuted with "scale matters". I commonly see arguments on all sorts of topics that make the same mistake you're calling out.

Re: AI is just unauthorised plagiarism at a bigger scale

#519
post #408

Earlier quoted context omitted.

Disney owns the 1992 production of Aladdin so who exactly are they "stealing" from?

I assume he's saying Disney owns the 1992 film so the 1999 film is not theft, but he wants it to be because he doesn't like the 1999 film. Thus the quotes.

That's not a charitable reading of the comment, and furthermore, it's not even a reasonable assumption. Other comments clarify that the "theft" is in quotes because it's a figurative theft, not from Disney to themselves, but from Disney to the earlier, non-copyrighted folk tales it drew inspiration from. And the "theft" is that the Disney IP supplanted (via ubiquity) the public domain versions to the point lots of people aren't even aware they exist. Nobody is arguing it's literal theft, hence the quotes.
Post reply on HN