Live data from Hacker News

AI-Shunning robots.txt

github.com

11–20 of 92 posts

Re: AI-Shunning robots.txt

#11
post #7

Nice. Let's all contribute to this...ideally, web-hosts should provide this sort of thing by default so we can starve AI companies from training data and combine it with other strategies to put them out of business for good.

How about AI from non-companies? Or genuinely non-profit or open projects? Also - out of curiosity - do you use any AI yourself?

> How about AI from non-companies? Or genuinely non-profit or open projects?

AI from any project will allow AI to be used commercially, and thus I oppose it. Moreover, I oppose AI on various other princincples even independent of this: it further isolates people and can be used to develop other technologies that are too powerful for us to handle. In short, I believe human beings en mass are too stupid to use AI.

> Also - out of curiosity - do you use any AI yourself?

I do not, or at least I try my best not too. In fact, I hate AI with a passion. Obviously, there may be products here and there that have used AI that I in turn use. What can you do? But I attempt to minimize any contact I have with AI: I don't use Grammarly, any form of auto-suggest, I use an ancient phone (and I RARELY use it, I hate smartphones), I don't use AI features in software such as AI-noise reduction, I turn off all automatic features in software that may have some AI behind it.

If I find out a website uses AI for content generation, I ban it and never visit again.

The other day I downloaded a text editor that looked cool but I deleted it because I realized it has an AI-console (even though I never used it).

I also work for a business and I convinced them not to use AI. We're an online magazine and it turns out the vast majority of our readers supported that decision.

In short, I am against AI because I believe it provides virtually no benefits to humanity, only detriments.

Re: AI-Shunning robots.txt

#12
post #8
post #4

I am curious, do we have any evidence that AI is adhering to robots.txt and isn’t ignoring it since they are not technically crawling in the traditional sense? Even if they are right now it would be a quick switch for them to just ignore it.

This is about crawling for training data by the look of things. Not sure if the CHatGPT browsing mode uses a different user-agent but most of the entries in that list look like crawlers.

I had assumed this is related to sites like chatgpt going out and searching with a specific request.

Regardless, my original question is still valid. The companies have already shown a lack of care about the data they train off of. So if ethics have already gone out the window, what is to stop them from ignoring this file if they are not already.

Re: AI-Shunning robots.txt

#14
Not that I'm arguing for or against preventing access from AI crawlers, but wouldn't it make more sense to block them at a higher level, e.g. the webserver, and not even give them the choice to obey/disobey robots.txt?

Re: AI-Shunning robots.txt

#16

Not that I'm arguing for or against preventing access from AI crawlers, but wouldn't it make more sense to block them at a higher level, e.g. the webserver, and not even give them the choice to obey/disobey robots.txt?

How would you propose doing so?

We could repurpose the evil bit.

Re: AI-Shunning robots.txt

#17
post #7

Earlier quoted context omitted.

How about AI from non-companies? Or genuinely non-profit or open projects? Also - out of curiosity - do you use any AI yourself?

> How about AI from non-companies? Or genuinely non-profit or open projects? AI from any project will allow AI to be used commercially, and thus I oppose it. Moreover, I oppose AI on various other princincples even independent of this: it further isolates people and can be used to develop other technologies that are too powerful for us to handle. In short, I believe human beings en mass are too stupid to use AI. > Al…

Likewise, I've unsubscribed from multiple paid Patreons and Substacks as soon as they started using AI to generate content. I'd rather see an amateur MSPaint scribble than some dall-e monstrosity at the head of a newsletter.

Re: AI-Shunning robots.txt

#18
post #9

The crawlers can simply stop identifying themselves via custom user agent, can't they? Also why are "AI" crawlers are worse than "normal" crawlers? Either way, this is an exercise in futility.

> Also why are "AI" crawlers are worse than "normal" crawlers?

A search engine will index your content to bring people to it through search. An AI crawler will take your content to recapitulate it and sell it to others. Obviously it's more complicated than this, but this is how one might see it who wishes to use this file.

> Either way, this is an exercise in futility.

Not necessarily disqualifying. Laws against theft are also futile, in the sense that honest people don't need them and dishonest people don't follow them, and history since at least Hammurabi has been replete with examples of such laws not stopping theft. And yet. Seems worth the calories it costs to say "for the record, I do not give my consent for what you're doing".

Re: AI-Shunning robots.txt

#19
post #10

This makes complete sense because, as we all know, AI companies are very concerned with respecting the rights of the people they steal data from, and totally won't just ignore this.

At least you show intent and can then potentially prove they are not respecting your wishes. It’s better than doing nothing.
Post reply on HN