Earlier quoted context omitted.
Hmm, I'm not sure you RTFA. He didn't search for the term "bing", he searched for the term "site:.bing.com/search". So it IS kinda a big deal that they're disrespecting the robots.txt and listing those pages anyway. I have also seen Google listing one of my domains for which I specifically disallowed all spiders (Bing doesn't show those domains FYI). My feeble attempt at separating my personal and professional person…
> So it IS kinda a big deal that they're disrespecting the robots.txt and listing those pages anyway. Did you RMFC? The robots.txt doesn't match what is indexed. As has been exhaustively pointed out in this thread by more than a few commenters, "Search" != "search". On that very search results page, we see #1 which is: > OLAC search - Bing OLAC is an unrelated site, and at one point it apparently linked to "www.bing.…
This is highly ridiculous. HTTP URLs are not mandated to be case-sensitive (though it's recommended), and clearly lots of sites use them in a case-insensitive way. Robots should either consider robots.txt in a case-insensitive way (even if I'm conscious that lot of them, including major ones, currently don't do that, which is precisely what I consider to be a problem, and this is supported by what happened here -- where google is risking to appear as a fool). The following article has perfectly good arguments in favor of case-insensitivity or even more clever handling: http://www.slicksurface.com/blog/2007-04/be-careful-robotstx...
Also; failing to respect conditions of use of a service because an automatic process is not safe enough is not a completely exonerating excuse for the operator of such process...