Live data from Hacker News

Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

news.ycombinator.com

41–50 of 296 posts

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#42
If AI needs explicit information and context, surely it should focus on improving its context recognition rather than trying to fix that by inserting even more training data.

Regardless, I do agree that something like a robots.txt for AI can be very useful. I'd like my website to be excluded from most AI projects and some kind of standardized way to communicate this preference would be nice, although I realize most AI projects don't exactly care about things like the wishes of authors, copyright, or ethical considerations. It's the idea that matters, really.

If I can use an ai.txt to convince the crawlers that my website contains illegal hardcore terrorist pornography to get it excluded from the datasets, that's another way to accomplish this I suppose.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#43

Using robots.txt as a model for anything doesn't work. All a robots.txt is is a polite request to please follow the rules in it, there is no "legal" agreement to follow those rules, only a moral imperative. Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. In the age of AI we need to better understand where copyright applies to it, and potentially need reform of copyright to ali…

if you do data mining in the EU you are legally required to respect robots.txt afaik

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#44
post #31
post #15

Earlier quoted context omitted.

> "We" now want stricter copyright law when it comes to AI, but at the same time shorter copyright duration... AI is being used to do copyright laundering, at the same time "we", the people who can't afford to run our own AI, are still subject to absurd rules that AI owners get to ignore, apparently.

The barrier to running an AI model is getting lower every day, so the threshold for ignoring copyright is getting lower with it.

You are mistaken if you think companies will allow common people to ignore copyright on their IP.

The only IP that will be allowed to be stolen is that of other common people.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#45
post #14

Using robots.txt as a model for anything doesn't work. All a robots.txt is is a polite request to please follow the rules in it, there is no "legal" agreement to follow those rules, only a moral imperative. Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. In the age of AI we need to better understand where copyright applies to it, and potentially need reform of copyright to ali…

> Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. Failing to solve every problem does not mean a solution is a failure. From sunscreen to seatbelts, the world is full of great solutions that occasionally fail due to statistics and large numbers.

> Failing to solve every problem does not mean a solution is a failure.

There is something to be said though to OP's point where it's actually better to do nothing than an AI.txt because it can give a false sense of security, which is obviously not what you want.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#46

Earlier quoted context omitted.

Ok, fair point, I may be being a little hyperbolic. But my point is that it's not a system that we should copy for preventing the use of content in training AI. It would become a useless distraction. If you "violate" a robots.txt the server administrator can choose to block your bot (if they can fingerprint it) or IP (if its static). With an ai.txt there is no potential downside to violating it - unless we get new le…

> But my point is that it's not a system that we should copy for preventing the use of content in training AI The purpose OP is suggesting in the submission is the opposite, help AI crawlers to understand what the page/website is about without actually having to infer the purpose from the content itself.

Isn't that the entire point of the semantic web?

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#47

Earlier quoted context omitted.

"Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare." I like the idea of "ai.txt" but those who eat resources rarely listen to ToS. Frankly, I serve 503s to all identifiable bots, unless they are on my explicit allow list.

Why not serve fake garbage indistinguishable from real content by a computer, like LLM output? Sending errors just incentivizes bot owners to fix the identifiable parts

> Sending errors just incentivizes bot owners to fix the identifiable parts

Nah. It'll just make them fake their identity so it is harder to tell the traffic is from a bot.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#48
post #14

Using robots.txt as a model for anything doesn't work. All a robots.txt is is a polite request to please follow the rules in it, there is no "legal" agreement to follow those rules, only a moral imperative. Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. In the age of AI we need to better understand where copyright applies to it, and potentially need reform of copyright to ali…

> Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. Failing to solve every problem does not mean a solution is a failure. From sunscreen to seatbelts, the world is full of great solutions that occasionally fail due to statistics and large numbers.

That's still not an argument to introduce ai.txt, because everything a hypothetical ai.txt could ever do is already done just as good (or not) by the robots.txt we have. If a training data crawler ignores robots.txt it won't bother checking for an ai.txt either.

And if you feel like rolling out the "welcome friend!" doormat to a particular training data crawler, you are free to dedicate as detailed a robots.txt block as you like to its user agent header of choice. No new conventions needed, everything is already on place.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#50

Using robots.txt as a model for anything doesn't work. All a robots.txt is is a polite request to please follow the rules in it, there is no "legal" agreement to follow those rules, only a moral imperative. Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. In the age of AI we need to better understand where copyright applies to it, and potentially need reform of copyright to ali…

The poster wants the opposite - a way of explicitly helping AI systems/etc to use their site. If people ignore it, they're just giving up a bit of help.
Post reply on HN