Had a conversation with a firm that wanted a distributed scraper built, and they really did not care about site usage policies. You would be fooling yourselves if you think such a firm cared about robots.txt or page tags. We warned them they would be sued eventually, to contact the site owners for legal access to the data, and issued a hard pass on the project. Probably they assumed if the indexing process was out of…
Robots.txt doesnt create a legal obligation. It’s just a set of rules saying “if you don’t follow these rules to politely crawl our site, we’ll block you from crawling our site”. Obviously “anything goes” in civil suits however - if someone is being absurdly egregious with their crawling there’s usually some exposure to one tort or another.
And Reddit has definitely become more proactive about scrapers. ;-)