Live data from Hacker News

AI is just unauthorised plagiarism at a bigger scale

axelk.ee

661–670 of 783 posts

Re: AI is just unauthorised plagiarism at a bigger scale

#661
post #536

Earlier quoted context omitted.

The problem is that the ad vendors couldn't keep it in their pants. The ads you're talking about are a common vector for delivering malware onto people's PCs, and absolutely destroy the usability of sites. Between tracking cookies, popups, full screen banners, autoplaying video, flashing ads, and their unbelievably high weight in bandwidth - the internet is fairly unusable if you don't block any ads Bear in mind that…

>The problem is that the ad vendors couldn't keep it in their pants. You might not know, many people don't, that ad vendors came to the table little over a decade ago to make a truce with Ad Block Plus. ABP and advendors both saw that an "ad supported internet" was unsupported with no ads. So ABP was looking to set terms for what would be deemed as acceptable ads. Creators/service providers get incentive, users get m…

>You might not know, many people don't, that ad vendors came to the table little over a decade ago to make a truce with Ad Block Plus. ABP and advendors both saw that an "ad supported internet" was unsupported with no ads. So ABP was looking to set terms for what would be deemed as acceptable ads. Creators/service providers get incentive, users get manageable ads.

I'm very aware of this, most ad vendors did not come to a truce with ad-block plus. ABP tried to position itself as the gatekeeper of what ads users were allowed to use (a hugely financially beneficial position for them), and immediately ended up letting through a bunch of terrible ads

It was a nice idea, but it was never going to work. There was simply too much money for the advertisers to make to allow abp to be the gatekeeper of ad content

The nature of ads has gotten significantly more invasive over time, and blocking ads today is a mandatory part of security. Ad companies *do not* have a god given right to track you, or infect your PC with malware

Users rioted because ABP did a terrible job at managing the situation

>What I have never seen though, and have zero examples of, is internet users trying to reconcile the situation. It's just a relentless entitlement to free everything, with a small fraction sometimes subscribing, and an even smaller fraction sometimes donating. The users are unquestionably the biggest assholes in this situation. They won't even acknowledge they have a problem.

As I mentioned in the comment you replied to, there are lots of alternative forms of advertising that users have not revolted against to anywhere near the same degree, eg sponsored content segments in youtube videos

Re: AI is just unauthorised plagiarism at a bigger scale

#663

There’s a fallacy that gets used a whole lot to justify things like this (not just with LLMs), and I see it in many of the comments here: If it’s OK (or at least negligible on a small scale), then it must be OK on a large scale. It usually goes something like: If I can make money by learning something from a web page, why does a computer making money by learning everything from everyone upset people so? It’s the same…

This is a great point. I think for coding, the wording of the MIT open source license makes it clear that copying and distributing the software is authorised on a small scale and it's very clear that the act of copying must involve a person. It provides distribution and modification rights to "any person obtaining a copy of the software" and explicitly requires attribution for any significant parts. Mass-ingesting th…

[dead]

Re: AI is just unauthorised plagiarism at a bigger scale

#664
post #529

Earlier quoted context omitted.

>when we are talking about a person learning something from a webpage or book and something completely different when a webpage or book is used to adjust some weights in a matrix What material differences exist between the two besides "humans good, computers bad"? >Calling that learning is a distraction from the real copyright violations going on. Most courts so far have ruled that it counts as fair use.

I thought fair use was dead after Napster

No, only the type of "fair use" that people slap on their youtube uploads, thinking them it gives them a "get out of jail free" card for copyright infringement. Fair use was repeatedly affirmed in the 2010s, eg. https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....

Re: AI is just unauthorised plagiarism at a bigger scale

#665
post #295

This site is strange. I'm pretty sure there's lots of AI shilling happening on it. I don't think the opinions here are authentic, they seem to be opinions that the AI company CEOs would hold, not the disenfranchised 99%. I used to trust HN, I'm not so sure I can now.

Any examples? There are obviously a lot of programmers here who think AI is a great tool and don't feel disenfranchised by it.

Get your LLM to find some for you.

Re: AI is just unauthorised plagiarism at a bigger scale

#666

There’s a fallacy that gets used a whole lot to justify things like this (not just with LLMs), and I see it in many of the comments here: If it’s OK (or at least negligible on a small scale), then it must be OK on a large scale. It usually goes something like: If I can make money by learning something from a web page, why does a computer making money by learning everything from everyone upset people so? It’s the same…

> It’s like if I go to Golden Gate Park and pick one flower, I shouldn’t do that, but no one cares. But if I build a machine to automatically cut every flower in the park because I want to sell them, that’s different. It's not like that, because flowers are a physical object and moving them to one place deprives their original location of the flowers. When an LLM learns something from a webpage, the webpage is still…

But you're still depriving the world of future flowers. Why spend years studying, sacrificing time with others, living frugally if others can take or monetize the result for free? Most people need compensation to justify their effort. Or the option to not have their years of work/sacrifice co-opted into an ai generated ad for toilet bowl cleaner.

No cost copying doesn't remove the need for compensation to sustain ongoing creation. Society has long treated knowledge, art, and thought as high-value outputs, and accepted the copyright tradeoff to support them. That is long settled and no 'get rid of copyright' proponents argue satisfactorily why the 300 year corpus of thought on that is invalid. Long copyright terms may justify reform but not rejection of the establishment that creative work needs economic value to sustain ongoing creation, and that ongoing creation is a net positive/desirable for society.

You are free to release copyright free today. In software that has unlocked immense value. In other areas those choosing copyright have unlocked more value. But software is different, I can get hired to build on the free. No one is hiring an author to expand their book to include fanfiction. And were that the model, it would arguably result in worse results as we are now back to the much worse patronage system where Bob hordes what he's paid for and only shares it with friends for status. For 300 years we've understood because of dynamics paywalled copyright with a throttled side of libraries unlocks the greatest access to knowledge. Eliminating duplication cost has not changed that.

'but I want every flower there is today and I don't care if there are any future flowers' doesn't change that, it's simply a new value judgement that my want/use case today outweighs the cost to society of lost future knowledge creation/return to a patronage based reward system. Again 300 years of thought say that results in a worse outcome for society. How does the typical OSS project that depends on patronage fare? Do we really want to return all knowledge output to that model?

Re: AI is just unauthorised plagiarism at a bigger scale

#667

There’s a fallacy that gets used a whole lot to justify things like this (not just with LLMs), and I see it in many of the comments here: If it’s OK (or at least negligible on a small scale), then it must be OK on a large scale. It usually goes something like: If I can make money by learning something from a web page, why does a computer making money by learning everything from everyone upset people so? It’s the same…

> It’s like if I go to Golden Gate Park and pick one flower, I shouldn’t do that, but no one cares. But if I build a machine to automatically cut every flower in the park because I want to sell them, that’s different. The problem here is, in your example the small scale example, and the large scale example are both unacceptable behavior. Learning from others at a small scale is not only socially acceptable, but is th…

> The problem here is, in your example the small scale example, and the large scale example are both unacceptable behavior.

no, not really, or at the very least they're not at all in the same category of "unacceptable behavior"

Re: AI is just unauthorised plagiarism at a bigger scale

#668

There’s a fallacy that gets used a whole lot to justify things like this (not just with LLMs), and I see it in many of the comments here: If it’s OK (or at least negligible on a small scale), then it must be OK on a large scale. It usually goes something like: If I can make money by learning something from a web page, why does a computer making money by learning everything from everyone upset people so? It’s the same…

If one person is murdered, that's bad. If a million people are murdered, that's war. If one word is stolen by AI, that's bad. If a million words are stolen by AI, that's business.

more like

If one word is stolen by Joe, that's bad. If a million words are stolen by Meta, that's business.

AI isn't the problem, is corporations using AI that are the problem

Re: AI is just unauthorised plagiarism at a bigger scale

#669

There’s a fallacy that gets used a whole lot to justify things like this (not just with LLMs), and I see it in many of the comments here: If it’s OK (or at least negligible on a small scale), then it must be OK on a large scale. It usually goes something like: If I can make money by learning something from a web page, why does a computer making money by learning everything from everyone upset people so? It’s the same…

Of course it's robbery. I don't think anyone is truly arguing it's not. The issue is that, if we don't do it, China will. Game over. I'm surprised I hvan't seen more economist scholars exploring this topic; it's a fastincating phenomenon. I've seen folks try and re-visit history and compare what's happening with AI to some historic event--but, we've never seen anything quite like it. As much as history repeats itself…

> The issue is that, if we don't do it, China will.

These AI companies aren’t state enterprises. How is geopolitics a justification?

If it were just the military training them, probably no one would care about the copyright infringement angle, it makes sense that the government could ignore those rules for national security.

But Mark Zuckerberg isn’t training his models to protect us from China. He’s doing it to make himself even more ridiculously wealthy.

Re: AI is just unauthorised plagiarism at a bigger scale

#670

Earlier quoted context omitted.

> It’s like if I go to Golden Gate Park and pick one flower, I shouldn’t do that, but no one cares. But if I build a machine to automatically cut every flower in the park because I want to sell them, that’s different. The problem here is, in your example the small scale example, and the large scale example are both unacceptable behavior. Learning from others at a small scale is not only socially acceptable, but is th…

> The problem here is, in your example the small scale example, and the large scale example are both unacceptable behavior. no, not really, or at the very least they're not at all in the same category of "unacceptable behavior"

The argument isn't small crime vs large crime. It is no crime regardless of scale.

If it is acceptable for a person to learn, then it should be acceptable for a machine. And any derived works produced from that information isn't theft or copyright violation.

Though I do think there is a valid gripe with the LLMs being trained on pirated materials. I've also personally learned from a lot of PDF of textbooks I didn't own.

Post reply on HN