Live data from Hacker News

Twitter just updated its robots.txt to exclude all scrapers

twitter.com

11–20 of 20 posts

Re: Twitter just updated its robots.txt to exclude all scrapers

#11
post #5

Earlier quoted context omitted.

Except don't they have a special deal with Google to use the firehose?

IMO, that's exactly the reason. Before search engines could scrape the data and load the content for free. Now they'll need to reach firehose data agreements.

Take a look at the robots file without the "www" subdomain. Its likely to prevent content duplication/push for only the non "www" url to appear on search engines.

https://twitter.com/robots.txt

Re: Twitter just updated its robots.txt to exclude all scrapers

#12
post #10

Earlier quoted context omitted.

Absolutely nothing from Twitter should be appearing in search engines.

Actually i believe its only from the "www" subdomain. Take a look at the robots without the "www" https://www.twitter.com/robots.txt https://twitter.com/robots.txt This is likely just to prevent content duplication/nudge users to visit without the "www".

yep makes a huge difference... as explained here as well http://webmarketingschool.com/no-twitter-did-not-just-de-ind...

Re: Twitter just updated its robots.txt to exclude all scrapers

#15
post #10

Earlier quoted context omitted.

Actually i believe its only from the "www" subdomain. Take a look at the robots without the "www" https://www.twitter.com/robots.txt https://twitter.com/robots.txt This is likely just to prevent content duplication/nudge users to visit without the "www".

yep makes a huge difference... as explained here as well http://webmarketingschool.com/no-twitter-did-not-just-de-ind...

You should never use Google's estimate as a real estimate, especially once it gets past 100. There are better ways to move content, like 301 redirects or rel=canonical.

Re: Twitter just updated its robots.txt to exclude all scrapers

#16

Here's a write up as to why they have made the changes. Was going to write it yesterday, but life got in the way: http://webmarketingschool.com/no-twitter-did-not-just-de-ind...

Why would you not use 301 redirects or just rel=canonical?

Re: Twitter just updated its robots.txt to exclude all scrapers

#17
post #16

Here's a write up as to why they have made the changes. Was going to write it yesterday, but life got in the way: http://webmarketingschool.com/no-twitter-did-not-just-de-ind...

Why would you not use 301 redirects or just rel=canonical?

There must be some platform issue is my best guess...

Re: Twitter just updated its robots.txt to exclude all scrapers

#18
post #11

Earlier quoted context omitted.

IMO, that's exactly the reason. Before search engines could scrape the data and load the content for free. Now they'll need to reach firehose data agreements.

Take a look at the robots file without the "www" subdomain. Its likely to prevent content duplication/push for only the non "www" url to appear on search engines. https://twitter.com/robots.txt

Uh... Excellent point. Seems the entire speculation is/was unfounded.

Re: Twitter just updated its robots.txt to exclude all scrapers

#19
post #16

Earlier quoted context omitted.

Why would you not use 301 redirects or just rel=canonical?

There must be some platform issue is my best guess...

They could at least have allowed the Internet Archive in the robots.txt, since the way things stand all www.twitter.com links will be unavailable from the Wayback Machine. That will obviously be a huge loss to researchers.

Re: Twitter just updated its robots.txt to exclude all scrapers

#20
post #19

Earlier quoted context omitted.

There must be some platform issue is my best guess...

They could at least have allowed the Internet Archive in the robots.txt, since the way things stand all www.twitter.com links will be unavailable from the Wayback Machine. That will obviously be a huge loss to researchers.

(Update: the Wayback Machine will be fine, using the twitter.com/robots.txt)
Post reply on HN