Live data from Hacker News

Facebook's robots.txt

facebook.com

11–20 of 23 posts

Re: Facebook's robots.txt

#13

Even Facebook's robots.txt has a hatred for my pseudo-anonymous browser settings. Facebook gives me this (for any page): "Sorry, something went wrong. We're working on getting this fixed as soon as we can."

robots.txt isn't enforced.

Maybe they should be. Gentleman's agreements do not apply to robots.

Re: Facebook's robots.txt

#16
post #10
post #8

So what does it mean by facebook whitelisting a scraping service? Do they actively block scrapers?

I could be wrong but I believe that the the default is that spiders are blocked and only the "User-Agents" listed are allowed to scrape (but not the disallow pages).

You are correct.

Re: Facebook's robots.txt

#18

Earlier quoted context omitted.

robots.txt isn't enforced.

Maybe they should be. Gentleman's agreements do not apply to robots.

And how exactly do you propose verifying that the user agent purporting to be Googlebot or Firefox is actually who they are? They're inherently unenforceable.

robots.txt is basically a list of rules that lay out "This is how we'd like you to crawl us. We might stop serving you if you don't comply", rather than a hard-and-fast set of directives that specify how a webcrawler will be guaranteed to behave.

Post reply on HN