Live data from Hacker News

Bing search results showing up in Google

jacquesmattheij.com

51–60 of 131 posts

Re: Bing search results showing up in Google

#51
post #23
post #22

Microsoft is way out of line. Google figures out what content is good by crawling every page and doing the leg work, and Bing copies Google data and displays what google displays. Google proved it with the bing sting. there is absolutly NO reason why bing should have linked to those documents, other than that they copied off of Google's exam paper. When students do this, it is called plagiarizing. The smoke getting t…

Google proved it with the bing sting They didn't prove anything. Every experiment needs a control. Where's Google's?

Google's hypothesis was "Does Bing use Google data in its rankings?" That hypothesis was proven (because there was no other way to get those crazy links except for clickstream data showing users going from a google search for those nonsense terms to those sites).

If you want to explore a very different hypothesis, namely "Does Bing single out Google in its weightings of clickstream data?" then I suggest you go here:

http://projectgus.com/files/googlebing/seaport-trace.txt

That's a packet capture of some clickstream data. That should be more than enough to forge as much data as you like. You can then make up a nonsense search term, like "doesb1ngtrustgmorethananyoneelse" that should get zero results, then forge clickstreams going from a google.com search to "yes.com" as well as an equal number of searches for that term going from some other search engine to "no.com" and you can explore the weights as much as you want.

That said, I'd say that the fact that they weight Google highly enough that they'd take their word for it that a clearly irrelevant term should be mapped to some site is strong evidence, I think.

Mind you, I don't think it's "wrong" exactly for Bing to do this. I'm not worried about it destroying search, either. The spammers/SEO types will make it useless soon enough.

Re: Bing search results showing up in Google

#52

Dear Microsoft, case-sensitivity is important. [1] m.bing.com/robots.txt says "/search", not "/Search". All of the crawled [2] urls are "/Search" or "/~/search". Also, wap.bing.com/robots.txt explicitly "Allow:"s several search pages, which are indexed by google. [1] EDIT: Case insensitivity is often important. Above comment notes that some robots are case-insensitive. I suspect Google is not, based on the results. […

Naturally, Bing is hosted on a Windows server, which inherits the Windows filesystem eccentricities. Case insensitivity is among those. Because "search" has six letters, that would mean that the robots.txt would need to have 64 entries to completely exclude this directory. That's not even including the tilde thing or any other paths to that directory. And that's for one directory.

Lame? Yes. Google's fault? Not in the slightest. But it brings up an interesting question: if MS clickstream gathering included an opt-out mechanism that happened to be impractical for Google, would that change the ethics of any of this? Say, by having the Bing toolbar identify itself in user-agent so that Google could block it if they wanted?

I wouldn't think that would materially change the situation. If Google really wanted to, they could probably "block" this now by encrypting their existing URL redirects, thus hiding the URL from the Bing toolbar entirely, at least until the user is out of the Google system.

Re: Bing search results showing up in Google

#53
post #22

Microsoft is way out of line. Google figures out what content is good by crawling every page and doing the leg work, and Bing copies Google data and displays what google displays. Google proved it with the bing sting. there is absolutly NO reason why bing should have linked to those documents, other than that they copied off of Google's exam paper. When students do this, it is called plagiarizing. The smoke getting t…

Yes, nothing exists on the internet if it's not on Google. When people click on things in Reddit, if Google hasn't yet indexed it, then it doesn't exist.

Re: Bing search results showing up in Google

#54
post #39

Earlier quoted context omitted.

Because by incorporating clickstream data from Google, they're effectively copying Google search results. Bing should blacklist Google from its clickstream data.

Let's say tomorrow DDG is the search engine with the largest market share. Then Bing would be getting all the clickstream data from DDG. I hope you do realize that this "algorithm" is not Google specific. Its just a novel ranking technique that incorporates a human user feedback loop and is a pretty well known technique in the information retrieval field.

It would be equally unethical to be copying DDG's results in this fashion.

Re: Bing search results showing up in Google

#56
post #28
post #7

This looks like microsoft is assuming robots.txt is case insensitive?

It is. Only domain names are case-insensitive. This document explains how Google handles robots.txt : http://code.google.com/web/controlcrawlindex/docs/robots_txt...

Are we reading the same file? Under where it describes matching paths (so replace /fish and /Fish with /search and /Search if you like):

==Example path matches==

[path] /fish

Matches: /fish /fish.html /fish/salmon.html /fishheads /fishheads/yummy.html /fish.php?id=anything

Does not match: /Fish.asp /catfish /?id=fish

Comments: Note the case-sensitive matching.

Re: Bing search results showing up in Google

#57
post #22

Microsoft is way out of line. Google figures out what content is good by crawling every page and doing the leg work, and Bing copies Google data and displays what google displays. Google proved it with the bing sting. there is absolutly NO reason why bing should have linked to those documents, other than that they copied off of Google's exam paper. When students do this, it is called plagiarizing. The smoke getting t…

They proved that Microsoft uses clickstream data to rank websites, and in 7% of manufactured cases that's all the data they have? I think this is exactly as petty and silly as last week's news, and now they get to spend a week explaining how this occurred and that they do obey robots.txt.

Everyone else in this thread has already explained that microsoft is not properly using robots.txt by having case insensitive urls, hence why these urls were indexed.

Re: Bing search results showing up in Google

#58
post #22

Microsoft is way out of line. Google figures out what content is good by crawling every page and doing the leg work, and Bing copies Google data and displays what google displays. Google proved it with the bing sting. there is absolutly NO reason why bing should have linked to those documents, other than that they copied off of Google's exam paper. When students do this, it is called plagiarizing. The smoke getting t…

> there is absolutly NO reason why bing should have linked to those documents There is absolutely a reason. A user queried for a string, then followed a link. Biasing Bing's search results towards the followed links is a signal that improves their search. > When students do this, it is called plagiarizing. The smoke getting thrown by MS is just to distract and divert while they scramble to hide what they did. When I…

There is absolutely a reason. A user queried for a string, then followed a link. Biasing Bing's search results towards the followed links is a signal that improves their search.

Yes, but the actual effect is that they've just copied Google's results, rather than extracted valuable insight from their own users. And I don't see how Bing could have scraped the search terms and results from a Google session unless they are using code specifically tailored for Google's site. Given that, the clickstream excuse rings awfully hollow.

Plagiarism is going to be a judgement call in a developing field like search. Google's judgement is that Bing's scraping of their site, in this case, is fairly underhanded. I would agree.

Re: Bing search results showing up in Google

#59

With respect to the author, the conclusion here is very flawed. If you search for Bing in Google, you get Bing all over page 1. If you search for Google in Bing, you get Google all over page 1. That's not the result of Google capturing click stream data from Google Chrome and copying Bing's results, nor is it the result of Microsoft capturing click stream data from IE8 and copying Google's results. That's just the na…

Except for the schema and host parts (which are not part of robots.txt anyway) URLs are case sensitive (ref: RFC 3986, sections 6.2.2.2 and 6.2.3).

The problem here is that Microsoft servers respond to /search, /Search and /SeaRCh without distinction. They are all distinct URLs. If it was the intended behavior (stupid, but understandable, coming from Microsoft), then robots.txt should contain all variants in capitalization for each path. A better solution would be to force a 301 redirect to a canonical path, and have this path in robots.txt. Google would work as expected.

The original article is totally bogus. I can't imagine how it has over 90 votes.

Re: Bing search results showing up in Google

#60

With respect to the author, the conclusion here is very flawed. If you search for Bing in Google, you get Bing all over page 1. If you search for Google in Bing, you get Google all over page 1. That's not the result of Google capturing click stream data from Google Chrome and copying Bing's results, nor is it the result of Microsoft capturing click stream data from IE8 and copying Google's results. That's just the na…

Hmm, I'm not sure you RTFA. He didn't search for the term "bing", he searched for the term "site:.bing.com/search". So it IS kinda a big deal that they're disrespecting the robots.txt and listing those pages anyway. I have also seen Google listing one of my domains for which I specifically disallowed all spiders (Bing doesn't show those domains FYI). My feeble attempt at separating my personal and professional person…

There's a bit of a weird myth with robots.txt and the idea that it prevents pages from showing up in search engine indexes. Robots.txt means that the search engine cannot crawl the page - it can still include it in the search results if it sees enough sites linking to it. It can take a guess at what the page title can be, but there's usually no descriptive snippet because it's unable to see what's on the page.

If you don't want the page to be included in the index at all, you can use the meta noindex tag, by putting this in the head:

Pro tip: you need to also unblock that page from robots.txt - if Google isn't allowed to crawl the page, it can't see the meta noindex tag, which means it would stay indexed.

Post reply on HN