Earlier quoted context omitted.
Archive.org mostly ignores robots.txt https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...
People always use that link as reference to say that Internet Archive ignores robots.txt but it only actually says they are ignoring it for government sites. It suggests that they might do it for other sites in the future (of 2017), but does not actually say that that they have done it. https://blog.archive.org/2018/04/24/addressing-recent-claims... which is a year later mentions that they have an automated process w…
If archive.org can be manipulated to remove content either via legal threats or simple robots.txt it loses a significant portion of its societal value.