Live data from Hacker News

AI is just unauthorised plagiarism at a bigger scale

axelk.ee

231–240 of 783 posts

Re: AI is just unauthorised plagiarism at a bigger scale

#231

Earlier quoted context omitted.

It's never been a problem with people ad-blocking for the last 20 years, why is it suddenly a problem now? We've been celebrating denying creators revenue for decades... Maybe this is just the internet hypocricy of "When I do it, it's good, when they do it, it's bad".

People usually point at the scale when this discussion comes up, in my experience. These companies are doing something at a huge scale spending tons of money to do it so the potential harm is greater. People can easily justify their own piracy because it’s small scale. Even when they organize, create a whole software and tooling ecosystem around pirating media to stick into jellyfin or plex. AI still did it bigger an…

Don't forget that the money being spent to do said scraping has, in great sums, come from subsidies paid by taxes from public coffers.

Re: AI is just unauthorised plagiarism at a bigger scale

#232

Earlier quoted context omitted.

It's never been a problem with people ad-blocking for the last 20 years, why is it suddenly a problem now? We've been celebrating denying creators revenue for decades... Maybe this is just the internet hypocricy of "When I do it, it's good, when they do it, it's bad".

People usually point at the scale when this discussion comes up, in my experience. These companies are doing something at a huge scale spending tons of money to do it so the potential harm is greater. People can easily justify their own piracy because it’s small scale. Even when they organize, create a whole software and tooling ecosystem around pirating media to stick into jellyfin or plex. AI still did it bigger an…

On the whole, about 35% of internet users are ad-blocking. In the tech space it's upwards of 70%.

It's in no way, shape, or form "small scale", and has fundamentally changed the the very nature of the internet for the worse (opinions/views of ad blocking people don't matter).

Re: AI is just unauthorised plagiarism at a bigger scale

#233

I don’t know if this author supports OSS but I’ll share this because HN generally is full of people with that mindset. It’s deeply ironic that if you forget about LLMs and look only at the outcome—-we’ve found a way to legally circumvent copyright and the siloing of coding knowledge, making it so you can build on top of (almost) the whole of human coding knowledge without needing to pay a rent or ask for permission—-…

That's not the reason why I publish OSS. I also publish that software under specific licenses that impose specific obligations (e.g., making the source available to users and attribution being given to the original author(s)).

Re: AI is just unauthorised plagiarism at a bigger scale

#234
post #171

if theres just one good thing coming out of ai its breaking copyright law forever. no one should be able to "own" ideas. royalties for commercial use is another thing and i support it but what we know as (non commercial) piracy and unlicensed fan art should be 100% legal

This is an incredibly naive view of intellectual property. If you cannot own things you create, there is little incentive to create and share those things. Do you think any of your favorite movies and TV shows ever get made without copyright protections? Of course not, because money needs to change hands for those things to be funded.

Re: AI is just unauthorised plagiarism at a bigger scale

#235
post #27

The broader problem of original sources not being given credit in a way that rewards them remains. Websites owners are paying to host their content so that spiders can come and crawl them and index it into the AI and then if they’re lucky, they might get a citation, but otherwise there’s very little reward for being a provider of content. And of course, this is something that’s getting worse and worse. Why look at a…

It's never been a problem with people ad-blocking for the last 20 years, why is it suddenly a problem now? We've been celebrating denying creators revenue for decades... Maybe this is just the internet hypocricy of "When I do it, it's good, when they do it, it's bad".

Total sleight of hand.

Ad blocking has always been a problem for creators but it's aimed at big corps - non-creators. The creators asked people to support them other ways or turn off the blocking. And it's not like the little independent creators wanted this version of commercialized internet in the first place.

The ai marketing teams are spinning everything they can but no AI companies are the conscript, the vultures. No question about it.

Re: AI is just unauthorised plagiarism at a bigger scale

#236
post #124

IP attorney here and actively working on this problem. nla: if you create content online (public repo code, blog, podcast, YouTube, publishing) the smartest thing you can do if to file a US copyright, even if you have a hobby blog. Anthropic paid $1.5B in a class settlement to authors because it was piracy of copyrighted works. If we as a HN community had our works protected, there are potentially huge statutory dama…

Wait what do you mean by "file a copyright"? I have never heard of this, all explanations of copyright I have heard say that you automatically own the copyright to the things you make; and that "all rights are reserved" by default unless you give up on them through granting a license. Is this no longer the case? Why is this now suddenly different? When did it change?

Briefly, there is default copyright and registered copyright. Registering works grants stronger protections (i.e. bigger fines if broken).

Re: AI is just unauthorised plagiarism at a bigger scale

#237
post #27

The broader problem of original sources not being given credit in a way that rewards them remains. Websites owners are paying to host their content so that spiders can come and crawl them and index it into the AI and then if they’re lucky, they might get a citation, but otherwise there’s very little reward for being a provider of content. And of course, this is something that’s getting worse and worse. Why look at a…

It's never been a problem with people ad-blocking for the last 20 years, why is it suddenly a problem now? We've been celebrating denying creators revenue for decades... Maybe this is just the internet hypocricy of "When I do it, it's good, when they do it, it's bad".

There is more to life than money.

Many of the websites I read do not collect any appreciable amount of money from ads, or have no ads at all (one example: news.ycombinator.com :) ). They want a recognition, or to share the knowledge, or community, or they are building their brand... And AI is destroying this all - the first result of "zx80" is an AI overview with a link to wikipedia and some youtube videos. If person stops there , they will never get to computinghistory.org.uk link, and won't see any related information about the variants and models.

Re: AI is just unauthorised plagiarism at a bigger scale

#238
post #58
post #27

The broader problem of original sources not being given credit in a way that rewards them remains. Websites owners are paying to host their content so that spiders can come and crawl them and index it into the AI and then if they’re lucky, they might get a citation, but otherwise there’s very little reward for being a provider of content. And of course, this is something that’s getting worse and worse. Why look at a…

About a year ago OpenAI crawled and go DDOS level the company I work. Even despite the robots.txt not allowing it, and despite some recaptcha we could assemble in time. We found our data in the outputs of their models but who can do anything about it...

> We found our data in the outputs of their models but who can do anything about it...

If the crawlers refuse to voluntarily respect your robots.txt, then you are well within your rights to poison their data.

Re: AI is just unauthorised plagiarism at a bigger scale

#239
At this point, I think google, openai, anthropic, etc already realise this and are just trying to pretend this isn't true. I even think some C-suite who are not in AI companies but are boosters know this too. This has been true since 2022 but they're hoping (likely correctly) that governments won't move fast enough to protect the IP of the actual productive class.

I think the long term reality is that the models still need training data so they fundamentally do need new writing/code/art to train on, and even then the usual issues like hallucination will still be with us. It's just the moment that actually hurts the (already questionable) profitability of the model peddlers, they will have gotten their IPOs and they can safely jump ship and the ultimate mess can be passed to the softbanks, the temaseks, and the governments of the world to clean up for them. What the future holds after the crash I'm not sure as the models won't disappear (especially now that the stolen data is already crystalised in open source models) but in the near term the mass theft that constitutes llms will become more and more understood even amongst the PMC and that in order to remain viable, you need the productive to keep producing, and unlike LLMs, you can't force them to do it without payment.

Post reply on HN