Live data from Hacker News

AI is just unauthorised plagiarism at a bigger scale

axelk.ee

501–510 of 783 posts

Re: AI is just unauthorised plagiarism at a bigger scale

#501
post #264

Earlier quoted context omitted.

robots.txt seems like it should be a legally-binding terms of service which would make them outright copyright infringing. Sue for $180,000 per infringement which should be calculated for each illegal API call.

Was your robots txt written by a lawyer? Does it hold up in the court?

Contracts are legally binding even if they weren't written by a lawyer. Copyright is legally binding even if no copyright claim is explicitly stated.

I looked into this a bit (not a lawyer) and it seems that robots.txt isn't legally binding to either party, but this seems to have two major implications for AI agents (and crawlers/scrapers in general).

First, even if the robots.txt says you can crawl the site, that isn't a copyright grant of any kind or permission to copy/use that data outside of the permissions granted by the TOS.

Second, ignoring the robots.txt while also pirating the site contents could point to bad-faith and makes a much stronger case for double-damage penalties due to willful infringement.

If the site TOS doesn't explicitly grant an AI agent rights to copy out the site content AND the AI agent is ignoring the robots.txt at the same time, it seems a lot more likely that there's a strong copyright infringement case against the agent owner.

Re: AI is just unauthorised plagiarism at a bigger scale

#502
post #264

Earlier quoted context omitted.

robots.txt seems like it should be a legally-binding terms of service which would make them outright copyright infringing. Sue for $180,000 per infringement which should be calculated for each illegal API call.

Was your robots txt written by a lawyer? Does it hold up in the court?

It doesn't have to be written by a lawyer. The robots.txt file is an administrative directive, by the webmaster of the website, that you, being a scraper, MUST NOT go to page x and/or y, or MUST NOT go to directory z. All the law would have to say is that it is a crime to not obey these directives. It's similar to trespassing: if I put a sign that says "DO NOT ENTER" in bright red letters on a door in my apartment, or "authorized people only!", that is still legally binding and a court isn't going to care that it wasn't lawyer-authored. The court will only care that you were told to not enter that area, but did so anyway.

Re: AI is just unauthorised plagiarism at a bigger scale

#503

There’s a fallacy that gets used a whole lot to justify things like this (not just with LLMs), and I see it in many of the comments here: If it’s OK (or at least negligible on a small scale), then it must be OK on a large scale. It usually goes something like: If I can make money by learning something from a web page, why does a computer making money by learning everything from everyone upset people so? It’s the same…

We ran into a lot of stuff like this in the early days of the web. For example, there was a lot of information that was "public" in that anyone could go to the city courthouse and ask to see the documents. But it changed in nature when you could suddenly look up anyone in the country by typing their name in your browser.

we used to ship mass lists of addresses and phone numbers to people in each town and it was fine/appreciated.

Re: AI is just unauthorised plagiarism at a bigger scale

#504
post #352

Earlier quoted context omitted.

How about requiring AI companies to pay creators for training rights? Alternatively, models trained on the commons must be owned by the commons. Right now these AI companies are trying to have it both ways: it’s The People’s Data for training on comrade but ownership is privatized.

Practically speaking, who is going to enforce such a regime? Do you really want to give Chinese companies such a huge competitive advantage, that they aren't subject to the same costs as western companies? How do you even sort out which "creators" are owed, and how much? It's next to impossible, and would drown the legal system in litigation; it would likely cause more problems than it solves. On top of which you can…

If models are trained on the collective whole, they must be owned by the collective whole. If you believe funding creators for the training of private models is too slow, inconvenient, or creates a global disadvantage, then embrace collective ownership.

Re: AI is just unauthorised plagiarism at a bigger scale

#506
post #503

Earlier quoted context omitted.

We ran into a lot of stuff like this in the early days of the web. For example, there was a lot of information that was "public" in that anyone could go to the city courthouse and ask to see the documents. But it changed in nature when you could suddenly look up anyone in the country by typing their name in your browser.

we used to ship mass lists of addresses and phone numbers to people in each town and it was fine/appreciated.

You ever had a bump in the night my guy?

Or a stalker?

Re: AI is just unauthorised plagiarism at a bigger scale

#507
post #504

Earlier quoted context omitted.

Practically speaking, who is going to enforce such a regime? Do you really want to give Chinese companies such a huge competitive advantage, that they aren't subject to the same costs as western companies? How do you even sort out which "creators" are owed, and how much? It's next to impossible, and would drown the legal system in litigation; it would likely cause more problems than it solves. On top of which you can…

If models are trained on the collective whole, they must be owned by the collective whole. If you believe funding creators for the training of private models is too slow, inconvenient, or creates a global disadvantage, then embrace collective ownership.

Sure, I wish everything was perfectly fair too. But how do you practically and REALISTICALLY proceed, and ensure you don't end up doing more damage than benefit? The road to hell is paved with good intentions. Everyone seems much more focused on complaining, and talking about "what's fair", than actually proposing concrete steps that would lead to a better world, without a significant risk of creating a worse one.

Re: AI is just unauthorised plagiarism at a bigger scale

#508
post #160

Earlier quoted context omitted.

Where was the coolness inflection point?

In the past three months I've shipped more code than I have in years. New php extension https://github.com/hparadiz/ext-gnu-grep A Demo showing how to stream webrtc to KDE Wayland overlay. https://github.com/hparadiz/camera-notif A fun little tool that captures stdout/stderr on any running process. https://github.com/hparadiz/bpf_write_monitor Then I upgraded my 10 year old hand written framework to a new version tha…

But where is the cool stuff?

Re: AI is just unauthorised plagiarism at a bigger scale

#509
post #46

Seriously how is this surprising? We all know AI companies stole troves of data to train their models, why do you think they'll stop? Have they faced consequences for the mass theft of copyrighted data? You can't steal or profit off of that data, but it's fine for them for whatever reason. I guess because they're a force for good in the world and are pushing humanity forward eh?

That data is not stolen. It's still there.

the income from the data, on the other hand...

Re: AI is just unauthorised plagiarism at a bigger scale

#510
Whether or not its technically copyright infringement isn't the main issue I have. Its mostly that it concentrates the ability to collect rent from all of the content in the world into the hands of the few corporations who can build data centers at scale. This is a huge problem. Why would I make a webpage, a news site, an online magazine, or create art commercially if it can be swept up into these models and cut me out of any incentive? If its not legally copyright infringement now we need a new legal framework around it because its an absolute tragedy for human creativity and small enterprise.
Post reply on HN