Live data from Hacker News

AI is just unauthorised plagiarism at a bigger scale

axelk.ee

461–470 of 783 posts

Re: AI is just unauthorised plagiarism at a bigger scale

#461
post #171

if theres just one good thing coming out of ai its breaking copyright law forever. no one should be able to "own" ideas. royalties for commercial use is another thing and i support it but what we know as (non commercial) piracy and unlicensed fan art should be 100% legal

The biggest problem is not the broken commercialization, but the broken attribution. People should be recognized, when they create art. Art is an important way of how we humans express ourselves.

If you penalize and stigmatize copying, you get broken attribution.

Re: AI is just unauthorised plagiarism at a bigger scale

#462

What do people imagine can be done about it at this point? Offer a concrete suggestion. Any law or tax against this will give a huge advantage to other countries. It's already over, there's no going back to a world where this didn't happen. Let's just hope some good comes of it.

> What do people imagine can be done about it at this point? Offer a concrete suggestion. Simple. Free the companies from copyright liability, but after X amount of time they are required to release everything into the commons. The weights, the training scripts and the full training data (appropriately processed so that it can only be used for training and not for people to easily pirate whatever works were used). Th…

I'm sympathetic, since I think copyright laws are far too extensive and generous. But it's not simple, there are a lot of companies that won't fall under your jurisdiction, and the question is if that will give them a competitive advantage that kills the industry for you, and ultimately costs you more than you gain.

Re: AI is just unauthorised plagiarism at a bigger scale

#465
post #148

Earlier quoted context omitted.

Is it possible able to host your website in a way so that it couldn't be found via search engines (and thus wouldn't be crawlable I hope)? I know this has repercussions on findability, but if that wasn't a concern, I'm curious how one might circumvent getting crawled.

Possible yes, probable not likely. The moment you're issued a certificate your domain will be shown in the Certificate Transparency logs which are constantly monitored from anyone who wants to find new sites.

....Yet another vector through which "security experts" has caused a waterbed problem. Let's secure the Internet, oh no! We made a centralized list of operating domains for hostile actors to guide attacks with!

Re: AI is just unauthorised plagiarism at a bigger scale

#466

There’s a fallacy that gets used a whole lot to justify things like this (not just with LLMs), and I see it in many of the comments here: If it’s OK (or at least negligible on a small scale), then it must be OK on a large scale. It usually goes something like: If I can make money by learning something from a web page, why does a computer making money by learning everything from everyone upset people so? It’s the same…

My complaint with your argument is that the word learn means one thing when we are talking about a person learning something from a webpage or book and something completely different when a webpage or book is used to adjust some weights in a matrix. Calling that learning is a distraction from the real copyright violations going on.

Re: AI is just unauthorised plagiarism at a bigger scale

#468

Earlier quoted context omitted.

You are not a developer so you don't understand you can compile to a binary without revealing your sources? No copyright -> No GPL -> anyone can release their own close source version of open source software. Why do you think GPL was create in the first place? We always had public domain you know.

My compilers work just fine? Perhaps I'm not sure what your point is.

My point is that you are unable to understand the difference between GPL and public domain.

Re: AI is just unauthorised plagiarism at a bigger scale

#469
post #453

Earlier quoted context omitted.

Worse, the constant AI scraping is actually costing content providers additional money for no return. At least Google/Bing/Yahoo scraping would then be used to provide links back to your content.

How do you distinguish Google/MS scraping for Gemini/Copilot vs Google Search/Bing? In the case of Google, the UA is the same and you are entirely at their mercy to honor the Google-Extended instructions in robots.txt Google has further complicated it with new search announcement blurring lines between regular search and AI search. And AI likes to not honor any licenses or instructions when it is hungry for training…

If company like Meta are downloading pirated books etc.. to train their AI, they will surely honor robots.txt.

Re: AI is just unauthorised plagiarism at a bigger scale

#470
post #442

There’s a fallacy that gets used a whole lot to justify things like this (not just with LLMs), and I see it in many of the comments here: If it’s OK (or at least negligible on a small scale), then it must be OK on a large scale. It usually goes something like: If I can make money by learning something from a web page, why does a computer making money by learning everything from everyone upset people so? It’s the same…

quantitative changes in an activity produce qualitative changes Well said!

It reminds me of a Stalin* quote: "Quantity has a quality all its own."

* Note that it may be misattributed to him

Post reply on HN