With the current situation you either assume that everything is not usable, or you just not care and crawl everything that you can reach.
Tell HN: We should start to add “ai.txt” as we do for “robots.txt”
61–70 of 296 posts
Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”
#62Using robots.txt as a model for anything doesn't work. All a robots.txt is is a polite request to please follow the rules in it, there is no "legal" agreement to follow those rules, only a moral imperative. Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. In the age of AI we need to better understand where copyright applies to it, and potentially need reform of copyright to ali…
In general without a fair use exemption or permission from robots.txt saving a copy of a website’s content to your own servers is copyright infringement.
Purely factual information like Amazon’s prices isn’t protected by copyright, but if you want to save artwork or source files to train AI, that’s a copyright issue even before you get into the possibility of your AI being considered a derivative work.
Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”
#63Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”
#64Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”
#65Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”
#66Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”
#67Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”
#68Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”
#69Why would anyone want ai to train on and monetize your content? If there was a way to block ai stealing content most people would opt to block it.