facebook's... https://www.facebook.com/robots.txt
That really blows my mind. I mean, how can they say that's any kind of "agreement"? I someone writes a curl/wget script wrapper & points it to the top 10 websites, they don't enter into any kind of written contract or agreement.
They're telling the public that it does not have permission to crawl the site which try have the right to do. What is the problem with that?