You don't scrape Facebook, Facebook scrapes you!
Facebook's robots.txt
11–20 of 23 posts
Re: Facebook's robots.txt
#12http://www.google.com/robots.txt
Re: Facebook's robots.txt
#13Even Facebook's robots.txt has a hatred for my pseudo-anonymous browser settings. Facebook gives me this (for any page): "Sorry, something went wrong. We're working on getting this fixed as soon as we can."
robots.txt isn't enforced.
Re: Facebook's robots.txt
#14http://disqus.com/humans.txt
Uh oh... Something didn't work. > http://disqus.com/human.txt
Re: Facebook's robots.txt
#15http://www.google.com/robots.txt
/* would have sufficed
Re: Facebook's robots.txt
#16Re: Facebook's robots.txt
#17http://disqus.com/humans.txt
Re: Facebook's robots.txt
#18Earlier quoted context omitted.
robots.txt isn't enforced.
Maybe they should be. Gentleman's agreements do not apply to robots.
robots.txt is basically a list of rules that lay out "This is how we'd like you to crawl us. We might stop serving you if you don't comply", rather than a hard-and-fast set of directives that specify how a webcrawler will be guaranteed to behave.
Re: Facebook's robots.txt
#19Re: Facebook's robots.txt
#20What is a User Agent: Yeti?