Earlier quoted context omitted.
Except don't they have a special deal with Google to use the firehose?
IMO, that's exactly the reason. Before search engines could scrape the data and load the content for free. Now they'll need to reach firehose data agreements.
Twitter just updated its robots.txt to exclude all scrapers
11–20 of 20 posts
Re: Twitter just updated its robots.txt to exclude all scrapers
#12Earlier quoted context omitted.
Absolutely nothing from Twitter should be appearing in search engines.
Actually i believe its only from the "www" subdomain. Take a look at the robots without the "www" https://www.twitter.com/robots.txt https://twitter.com/robots.txt This is likely just to prevent content duplication/nudge users to visit without the "www".
Re: Twitter just updated its robots.txt to exclude all scrapers
#13Nah. https://twitter.com/robots.txt They blocked robots on their marketing pages.
Re: Twitter just updated its robots.txt to exclude all scrapers
#14http://webmarketingschool.com/no-twitter-did-not-just-de-ind...
Re: Twitter just updated its robots.txt to exclude all scrapers
#15Earlier quoted context omitted.
Actually i believe its only from the "www" subdomain. Take a look at the robots without the "www" https://www.twitter.com/robots.txt https://twitter.com/robots.txt This is likely just to prevent content duplication/nudge users to visit without the "www".
yep makes a huge difference... as explained here as well http://webmarketingschool.com/no-twitter-did-not-just-de-ind...
Re: Twitter just updated its robots.txt to exclude all scrapers
#16Here's a write up as to why they have made the changes. Was going to write it yesterday, but life got in the way: http://webmarketingschool.com/no-twitter-did-not-just-de-ind...
Re: Twitter just updated its robots.txt to exclude all scrapers
#17Here's a write up as to why they have made the changes. Was going to write it yesterday, but life got in the way: http://webmarketingschool.com/no-twitter-did-not-just-de-ind...
Why would you not use 301 redirects or just rel=canonical?
Re: Twitter just updated its robots.txt to exclude all scrapers
#18Earlier quoted context omitted.
IMO, that's exactly the reason. Before search engines could scrape the data and load the content for free. Now they'll need to reach firehose data agreements.
Take a look at the robots file without the "www" subdomain. Its likely to prevent content duplication/push for only the non "www" url to appear on search engines. https://twitter.com/robots.txt
Re: Twitter just updated its robots.txt to exclude all scrapers
#19Earlier quoted context omitted.
Why would you not use 301 redirects or just rel=canonical?
There must be some platform issue is my best guess...
Re: Twitter just updated its robots.txt to exclude all scrapers
#20Earlier quoted context omitted.
There must be some platform issue is my best guess...
They could at least have allowed the Internet Archive in the robots.txt, since the way things stand all www.twitter.com links will be unavailable from the Wayback Machine. That will obviously be a huge loss to researchers.