Live data from Hacker News

Matt Cutt's thoughts on Google Bing Debate

mattcutts.com

131–140 of 181 posts

Re: Matt Cutt's thoughts on Google Bing Debate

#131
post #121

Earlier quoted context omitted.

Of course not. Try inquiring google or bing about obtaining all the information they have gathered from your toolbar and see what response you get. Actually I'll tell you: "The gathered data is anonymous and we have no way of verifying who sent it to us". So how exactly do you own information that you have absolutely no access to and with no way to transfer to competing search companies?

That’s decidedly not what owning your clickstream data means. The argument is that users should be able to decide to give away their data, whether they can later access that data is irrelevant. Access to the collected clickstream data is simply not part of the agreement and as long as users are not coerced or tricked into agreeing that’ certainly unproblematic. (I can similarly agree to give away my photo collection…

You just said you can give away your photos by agreeing to not have any access rights to it and still own those same photos. It seems like you are using a different definition of "ownership" or maybe "click stream" than I am so this argument is pointless unless you define exactly what you mean by those words. By "click stream" I mean the aggregate history of clicks along with contextual information related to those clicks as one coherent, single entity and not individual clicks looked at one by one each of which was owned by the person that generated the click and who transfered ownership of that single click to somebody else. If you meant giving away a copy of your collection then that's a little different but that's not what's happening with user generated click streams. The only click stream copy is the one at google, microsoft headquarters and you have no access to it so you don't technically own it.

Re: Matt Cutt's thoughts on Google Bing Debate

#132

Earlier quoted context omitted.

Imagine I launch a search engine with no data. Then I feed it with urls IE users click on after their google search. I will eventually end up with an exact copy of google database. So I think that this technique can be called "copying". Now if Bing uses this technique for 0.1% of their data, then it can be said that 0.1% of their data are copied from Google database.

That's a straw man argument. That's not what they did. They had a search engine with data (lots of it), using many of the same factors that Google does (page/domain authority, on-page markup, etc). The feed that algorithm a lot of data. One bit of data they feed it is search behvior of toolbar users (presumably ALL search engines). In your scenario, that'd be a copy. In the reality scenario (described above), I think…

Ironically, you are refuting as a straw man a proposition I never made. I have never claimed Bing did that. I just highlight the fact the way they get data from google is technically a copy, since it has the capacity to duplicate the entire google database (which again is not what they are doing).

Re: Matt Cutt's thoughts on Google Bing Debate

#133

Earlier quoted context omitted.

> Google provides one piece: 1) An ordered list That's not an accurate view of how search works. The user provides the query, but the results that are returned don't need to have anything obvious to do with the query. The terms might not be on the page, and might not be found in any data remotely associated with the page. We don't just intersect posting lists here. Furthermore, with suggest and instant, often the que…

I didn't say the ordered list was obvious. But it is an ordered list (at least to the user -- it may not be ordered all the way down or partially so). Fair enough point about suggest and instant. In those cases Google is adding value to the query, but I think we'd both agree that it is incrementally so. Instant generally just helps you get to the eventual query faster. Spell checking does add value, but its a +epsilo…

>The association between the query and the URL is the least important aspect.

I don't understand why you believe this. This is the entire basis for why search is hard. It is the hard problem that search tries to solve. Many companies have spent in aggregate, billions of dollars in R&D trying to solve this, and all but a few have folded. It's an "AI-complete" problem in that solving it perfectly would be a sufficient demonstration of strong AI. It's the whole reason why Google exists.

Re: Matt Cutt's thoughts on Google Bing Debate

#134

Earlier quoted context omitted.

"Google created an artificial scenario where the ONLY input was Google search behavior and lo, the search results are exactly the same." Cutts addresses this: "As we said in our blog post, the whole reason we ran this test was because we thought this practice was happening for lots and lots of different queries, not simply rare queries."

If Google think they have other evidence, let's see it. The burden of proof is on the accuser, I think. (c/p from other thread.)

If this isn't enough to convince you, nothing will be. I'm not sure how the evidence could be any more convincing.

Content from Google's result database ends up used to generate Bing's results, full stop. Proved. It may be technical error, it may be only used a little bit, it may be deliberate, it may be accidental and perfectly ethical by some people's standards, there's all kinds of ways to legitimately interpret this fact. But Google has demonstrated beyond a shadow of a doubt that input from Google is used to generate some non-zero bit of Bing's results. There is absolutely, positively no other explanation that accounts for the facts of the case.

Re: Matt Cutt's thoughts on Google Bing Debate

#136
post #121

Earlier quoted context omitted.

That’s decidedly not what owning your clickstream data means. The argument is that users should be able to decide to give away their data, whether they can later access that data is irrelevant. Access to the collected clickstream data is simply not part of the agreement and as long as users are not coerced or tricked into agreeing that’ certainly unproblematic. (I can similarly agree to give away my photo collection…

You just said you can give away your photos by agreeing to not have any access rights to it and still own those same photos. It seems like you are using a different definition of "ownership" or maybe "click stream" than I am so this argument is pointless unless you define exactly what you mean by those words. By "click stream" I mean the aggregate history of clicks along with contextual information related to those c…

No. You can give something you own away. You obviously don’t own it anymore after you gave it away. Being able to give what you own away is part of ownership. (You can’t give something away you don’t own.)

Since you can still make a local copy of your clickstream data (every browser has some sort of history function) you don’t even really give it away, you keep it and hand over a copy (maybe not technically but it is still the exact same data) to Bing.

Here is an analogy: Imagine a scale which, after you agreed to it, automatically sends your weight to the manufacturer every time you use it. The manufacturer made it clear when you agreed that you will not be able to access your weight history. While slightly creepy, I don’t see any ethical or legal problems with that. You own the weight displayed on the scale, you can therefore also agree to send this weight to whomever you want, even if they don’t let you access your weight history. You are also free to make your own local copy of your weight history (for example by writing you weight on a piece of paper every time you use the scale).

Re: Matt Cutt's thoughts on Google Bing Debate

#137
Most of my thoughts have already been echoed elsewhere in this thread, but:

> "To me, what the experiment proved was that clicks on Google are being incorporated in Bing’s rankings."

It proved that clicks on other websites are being incorporated in Bing's rankings, which had already been public knowledge, I think. It didn't prove that only or disproportionately clicks on Google are being thus incorporated, although that is what Google is repeatedly claiming.

> "If clicks on Google really account for only 1/1000th (or some other trivial fraction) of Microsoft’s relevancy, why not just stop using those clicks and reduce the negative coverage and perception of this?"

Is Matt Cutts suggesting that Bing special-case an exclusion for Google results?

Re: Matt Cutt's thoughts on Google Bing Debate

#138
That video was way more interesting than this semi-fabricated drama. Search result spam, adsense role in it, and how to (not) tackle it is what I'd like to hear more about. Google search results degraded over the years to the point where I have a bing search on my wp7 phone as default search and I actually don't mind/don't care, because search results from google are not an imperative for me anymore (and I'm not cheerleading for anyone).

Re: Matt Cutt's thoughts on Google Bing Debate

#139

Earlier quoted context omitted.

I didn't say the ordered list was obvious. But it is an ordered list (at least to the user -- it may not be ordered all the way down or partially so). Fair enough point about suggest and instant. In those cases Google is adding value to the query, but I think we'd both agree that it is incrementally so. Instant generally just helps you get to the eventual query faster. Spell checking does add value, but its a +epsilo…

> The association between the query and the URL is the least important aspect. I don't understand why you believe this. This is the entire basis for why search is hard. It is the hard problem that search tries to solve. Many companies have spent in aggregate, billions of dollars in R&D trying to solve this, and all but a few have folded. It's an "AI-complete" problem in that solving it perfectly would be a sufficient…

The reason is that with the URL and the query, there's a decent chance you can derive the association.

Now the reason I say this is that to my best approximation the key difference that makes Google better than Bing (when it is) is the index size, and not the search algorithm.

So the two key pieces of value are: A) queries that people don't do on your site, which would have generated poor relevance

B) Pages that are actually browsed that you don't have indexed or up to date.

Now don't get me wrong. There is value in the association, but I think I'd capture 80% of the value with the above. Now as you point out, these lists are often non-obvious, but Bing does a near equal job of creating the lists. And if you give them the search terms where they have gaps. And fill the index, then I think the gap closes most of the way.

And lets be clear, because they clicked the link doesn't mean its a good link. See ehow.com or expertsexchange. But it does have some value.

And frankly I think MS would have been willing to let it go if Google had gone straight to MS with it and said, "We have this data. Even though its not unethical we can spin it to the media to make it look bad. Kill the association." MS probably would have as the net win isn't that huge. Now the PR loss in pulling it would be worse than the PR loss in keeping it (with no integrity loss, since they [and I] think it is perfectly ethical).

Re: Matt Cutt's thoughts on Google Bing Debate

#140
post #5

I think Matt lays out a strong argument for why they have made such a stink. And it sounds like a pretty compelling defense for their actions to go public. But... 1) I think they will regret it bigtime when all the attention they are causing makes the bored government officials poke their head in and realize that at a macro level, neither Bing nor Google does anything to protect user search privacy. 2) I think Google…

> They should focus on trying to be innovative again, that was the Google I respected. I work in search quality at Google. I'm busting my ass every day working on fundamental reimaginings of how results get ranked. I'm going to keep doing that regardless of what bing does, because it makes the world a better place, and it's fun. But, suppose the stuff I'm working on works out, and tomorrow Google shows up with a whol…

If they don't understand the results and how they got there, how will they improve them?
Post reply on HN