Live data from Hacker News

Matt Cutt's thoughts on Google Bing Debate

mattcutts.com

141–150 of 181 posts

Re: Matt Cutt's thoughts on Google Bing Debate

#141
I'm shocked at the amount of suppport on HN (thus far) for Bing's attitude on this issue.

I understand that they may well not target Google's SERPs specifically in their clickstream analysis but they should certainly have excluded Google from it, for ethical reasons.

Google state that Bing created associations from clickstreams through Google's SERPs on common queries (e.g. the tarsorrhaphy spell check test), not just long tail queries. Given that Google is extremely popular, this must have given a lot of weight to clickstream signals resulting from Google SERPs on many occasions, for common queries. That is entirely unethical and I'm shocked so many here don't find a problem with this.

Take the case of a highly-ranked great result on Google for a particular term, which Bing rates lower due to inferior algorithms. Bing's analysis of Google users would send that result higher in the Bing SERPs, mainly due to Google's expertise in highlighting that site, and only to a small degree due to the user's choice of clicking on it. That to be does fall under "copying" Google's results, it may not be illegal, or intentional, but copying it remains.

Bing should have excluded Google from their clickstreams, and I certainly hope Google exclude Bing from their's. (Matt Cutts stated they do in the video.)

Re: Matt Cutt's thoughts on Google Bing Debate

#142
post #117

Earlier quoted context omitted.

Yup this is the issue Google wants to raise. Expect Google lobbying soon for some law along the lines of "clicks on somesite.com belong to somesite.com and can only be shared with another party with somesite.com's permission.". Then in some Google ToS the "user" will be granted access to the clicks that the user makes on Google's search results page... but third party apps like Bing toolbar won't have access.

A far-fetched hypothetical scenario that has little basis in fact.

Perhaps :-) Google is a very smart, calculating company. They wouldn't dedicate so many resources to tricking the Bing toolbar if they weren't interested in getting something in return. Perhaps what Google wants in return is avoiding the following scenario: MS bundles the Bing toolbar with IE and data sharing is default opt-in. MS gets a ton of clickstream data from people using Google. Bing takes some of Google's marketshare.

Re: Matt Cutt's thoughts on Google Bing Debate

#143

Most of my thoughts have already been echoed elsewhere in this thread, but: > "To me, what the experiment proved was that clicks on Google are being incorporated in Bing’s rankings." It proved that clicks on other websites are being incorporated in Bing's rankings, which had already been public knowledge, I think. It didn't prove that only or disproportionately clicks on Google are being thus incorporated, although t…

Special-case an exclusion for any site that contains robots.txt?

Re: Matt Cutt's thoughts on Google Bing Debate

#144
post #44

I still don't understand why Google hasn't done the exact same test with a control variable. Run the exact same test again on a domain that isn't google.com.

Do we know they haven't? We only know they haven't reported on such a test.

Why would they purposefully make their case inconclusive?

Re: Matt Cutt's thoughts on Google Bing Debate

#145
post #126

Earlier quoted context omitted.

If Google's terms of use say you cannot share their search results with anyone else, are you legally allowed to send the clickstream to Microsoft? So if my dad calls with a question, and I search for the answer, I can't send him the link that I got from Google? Or do you mean something else?

That is exactly what I mean. IF Google say you cannot share their data with anyone else, then you are not legally allowed to share the answer with your Dad. IF Google say you must send a 'thank you' email to them after each query, you are legally obliged to do that. IF Google say you cannot share their data in an automatic, 'clickstream' fashion, then you cannot legally use Microsoft's toolbar. And if you break that…

I think Google would require a clickthrough TOS to enforce this. AFAIK Google has no TOS listed on its page, but they do have a copyright notice. And a link would not be sufficient.

But yes, if they did a clickthrough TOS then they could enforce that. Of course the day Google adds a clickthrough TOS there will be some serious scrutinizing.

Re: Matt Cutt's thoughts on Google Bing Debate

#146

Earlier quoted context omitted.

> The association between the query and the URL is the least important aspect. I don't understand why you believe this. This is the entire basis for why search is hard. It is the hard problem that search tries to solve. Many companies have spent in aggregate, billions of dollars in R&D trying to solve this, and all but a few have folded. It's an "AI-complete" problem in that solving it perfectly would be a sufficient…

The reason is that with the URL and the query, there's a decent chance you can derive the association. Now the reason I say this is that to my best approximation the key difference that makes Google better than Bing (when it is) is the index size, and not the search algorithm. So the two key pieces of value are: A) queries that people don't do on your site, which would have generated poor relevance B) Pages that are…

>is the index size, and not the search algorithm

I'm pretty thoroughly sure this is false. I'm sorry I can't show data to back it up, but Google has many many systems that are years ahead of bing's technology. I'm kinda hamstrung here by being unable to reveal anything about them. To a decent approximation, both Google and Bing likely have the whole internet that matters (and that they are allowed to crawl) in their indexes. Bing is just unable to return this data for as many queries as Google is.

>but Bing does a near equal job of creating the lists.

To the extent that this is true, how much of it is due to data harvested from Google search? This is something that google can't really demonstrate with evidence, and what I would have hoped that bing would clarify in a public statement if they had anything defensible to say.

Re: Matt Cutt's thoughts on Google Bing Debate

#147
post #87

Earlier quoted context omitted.

robots.txt is a good example of something that was introduced to make the new search engine technology more ethical but I'm not sure what point you are trying to make. If you're saying that we need a new robots.txt option saying "don't use clickstream data to help people find this page", I don't disagree; I'm not sure how many sites would take advantage of it though.

You're addressing the letter of the "law" not the spirit. Do we really need a new formal standard to indicate that this is unethical? To me it's pretty plain from the standards the web has already agreed upon.

To others it's just as plain that Microsoft's behavior is ethical. I'm still waiting to hear more before making up my mind. I don't blame you for being upset and frustrated, and can see why you think it's unethical. But a lot of people see it differently.

It would be great for Google and Microsoft to both come clean about all the different factors they use in their search but I'm not holding my breath.

Re: Matt Cutt's thoughts on Google Bing Debate

#148

Earlier quoted context omitted.

The reason is that with the URL and the query, there's a decent chance you can derive the association. Now the reason I say this is that to my best approximation the key difference that makes Google better than Bing (when it is) is the index size, and not the search algorithm. So the two key pieces of value are: A) queries that people don't do on your site, which would have generated poor relevance B) Pages that are…

>is the index size, and not the search algorithm I'm pretty thoroughly sure this is false. I'm sorry I can't show data to back it up, but Google has many many systems that are years ahead of bing's technology. I'm kinda hamstrung here by being unable to reveal anything about them. To a decent approximation, both Google and Bing likely have the whole internet that matters (and that they are allowed to crawl) in their…

To a decent approximation, both Google and Bing likely have the whole internet that matters (and that they are allowed to crawl) in their indexes. Bing is just unable to return this data for as many queries as Google is.

Maybe this is true, but its not apparent. I had commented on this a month or so back that I thought Bing was better at "vague" queries, where I don't know exactly what I'm looking for. But I'll know it when I see it.

Whereas Google is really good at targeted queries. I need info on the HP battery model number 003D434F90. These searches in Bing will often bring back literally nothing, while Google will often have one or two links, but they happen to be the link that I want. The text of the query is almost always found in these pages.

From that I infer it is index size, since the text is in the page.

To the extent that this is true, how much of it is due to data harvested from Google search?

While I find the quality of searches similar I don't find the results to be similar, if that makes sense. If Bing is harvesting your results, they're still using other very clever methods to surface other equally good, yet different results.

Re: Matt Cutt's thoughts on Google Bing Debate

#149
post #5

I think Matt lays out a strong argument for why they have made such a stink. And it sounds like a pretty compelling defense for their actions to go public. But... 1) I think they will regret it bigtime when all the attention they are causing makes the bored government officials poke their head in and realize that at a macro level, neither Bing nor Google does anything to protect user search privacy. 2) I think Google…

> They should focus on trying to be innovative again, that was the Google I respected. I work in search quality at Google. I'm busting my ass every day working on fundamental reimaginings of how results get ranked. I'm going to keep doing that regardless of what bing does, because it makes the world a better place, and it's fun. But, suppose the stuff I'm working on works out, and tomorrow Google shows up with a whol…

You have computed a relationship between FOO and BAR. For as long as you keep it to yourself, this ephemeral relationship is a secret, yours to keep.

Once you tell the world that FOO and BAR are related, and the whole world looks at it and says "Yay, it IS related!", the relationship stops being ephemeral and becomes actual, reflected in actions of the users. You no loner own that relationship.

You can't claim ownership of facts, even if you discovered them first, or created them into existence.

Re: Matt Cutt's thoughts on Google Bing Debate

#150

"Copying" was a pretty brutal word to use-- not surprising that it raised MSFT's hackles a bit. MS clearly uses toolbar users' clickstreams (on and off Google) to improve their own search efforts. Google created an artificial scenario where the ONLY input was Google search behavior and lo, the search results are exactly the same. Whether or not that steps over a line (I don't feel that it does), it's not "copying" in…

Imagine I launch a search engine with no data. Then I feed it with urls IE users click on after their google search. I will eventually end up with an exact copy of google database. So I think that this technique can be called "copying". Now if Bing uses this technique for 0.1% of their data, then it can be said that 0.1% of their data are copied from Google database.

Stop using the c... word. Call it user behavior analysis side effect.
Post reply on HN