Twitter just updated its robots.txt to exclude all scrapers
1–10 of 20 posts
Re: Twitter just updated its robots.txt to exclude all scrapers
#2Same file as of a few hours ago:
https://web.archive.org/web/20150715164726/https://twitter.c...
Re: Twitter just updated its robots.txt to exclude all scrapers
#3so what does that mean?
Re: Twitter just updated its robots.txt to exclude all scrapers
#4so what does that mean?
Absolutely nothing from Twitter should be appearing in search engines.
Re: Twitter just updated its robots.txt to exclude all scrapers
#5Re: Twitter just updated its robots.txt to exclude all scrapers
#6Re: Twitter just updated its robots.txt to exclude all scrapers
#7Earlier quoted context omitted.
Absolutely nothing from Twitter should be appearing in search engines.
Except don't they have a special deal with Google to use the firehose?
IMO, that's exactly the reason. Before search engines could scrape the data and load the content for free. Now they'll need to reach firehose data agreements.
Re: Twitter just updated its robots.txt to exclude all scrapers
#8Nope. No. They didn't.
What they did was some perfectly legitimate duplicate content protection.
Will write it up in a bit more detail...
Re: Twitter just updated its robots.txt to exclude all scrapers
#9Nah.
https://twitter.com/robots.txt
They blocked robots on their marketing pages.
Re: Twitter just updated its robots.txt to exclude all scrapers
#10so what does that mean?
Absolutely nothing from Twitter should be appearing in search engines.
Actually i believe its only from the "www" subdomain. Take a look at the robots without the "www"
https://www.twitter.com/robots.txt https://twitter.com/robots.txt
This is likely just to prevent content duplication/nudge users to visit without the "www".