Live data from Hacker News

Twitter just updated its robots.txt to exclude all scrapers

twitter.com

1–10 of 20 posts

Re: Twitter just updated its robots.txt to exclude all scrapers

#7
post #5

Earlier quoted context omitted.

Absolutely nothing from Twitter should be appearing in search engines.

Except don't they have a special deal with Google to use the firehose?

IMO, that's exactly the reason. Before search engines could scrape the data and load the content for free. Now they'll need to reach firehose data agreements.

Re: Twitter just updated its robots.txt to exclude all scrapers

#10
post #3

so what does that mean?

Absolutely nothing from Twitter should be appearing in search engines.

Actually i believe its only from the "www" subdomain. Take a look at the robots without the "www"

https://www.twitter.com/robots.txt https://twitter.com/robots.txt

This is likely just to prevent content duplication/nudge users to visit without the "www".

Post reply on HN