They do have a robots.txt [1] that disallows robot access to the spigot tree (as expected), but removing the /spigot/ part from the URL seems to still lead to Spigot. [2] The /~auj namespace is not disallowed in robots.txt, so even well-intentioned crawlers, if they somehow end up there, can get stuck in the infinite page zoo. That's not very nice. [1]: https://www.ty-penguin.org.uk/robots.txt [2]: https://www.ty-pen…
So? What duty do web site operators have to be "nice" to people scraping your website?