Live data from Hacker News

Matt Cutt's thoughts on Google Bing Debate

mattcutts.com

31–40 of 181 posts

Re: Matt Cutt's thoughts on Google Bing Debate

#31
"Copying" was a pretty brutal word to use-- not surprising that it raised MSFT's hackles a bit.

MS clearly uses toolbar users' clickstreams (on and off Google) to improve their own search efforts. Google created an artificial scenario where the ONLY input was Google search behavior and lo, the search results are exactly the same. Whether or not that steps over a line (I don't feel that it does), it's not "copying" in my book.

Another interesting point is that Google has been beating the "open" drum for a long time now. No walled gardens, right? If a Facebook user should have the right to take his data with him wherever he wants to go, shouldn't a Google user be able to fork over their behavior data to Bing?

Matt's point about MS' lack of clarity when getting folks permission to grab their clickstream was dead-on. THAT is pretty outrageous and MS should be ashamed of that.

Regardless of all that, hats off to Matt for keeping a cool head and stating his position in a respectful way.

Re: Matt Cutt's thoughts on Google Bing Debate

#32

Let’s take that thought to its conclusion. If clicks on Google really account for only 1/1000th (or some other trivial fraction) of Microsoft’s relevancy, why not just stop using those clicks and reduce the negative coverage and perception of this? And if Microsoft is unwilling to stop incorporating Google’s clicks in Bing’s rankings, doesn’t that argue that Google’s clicks account for much more than 1/1000th of Bing…

I wonder if you would be ok if Bing did the same thing to amazon. That is, imagine they used toolbar/IE logs to infer that people went to amazon, searched for LCD TVs and then purchased model X. Then they could boost pages about X in search, or implement a "bestselling" feature. After all, they are "just" using clickstream data here. Similarly, they could track Netflix, and so on. IMHO, there's a loophole in Bing's a…

You're right. There seems to be a gray area around whom owns the "click data":

  - "Are users entitled to give away for free their navigation history ?"
  - "To what extend are search engines allowed to implement specificities based on that data ?"

Re: Matt Cutt's thoughts on Google Bing Debate

#33

Earlier quoted context omitted.

I wonder if you would be ok if Bing did the same thing to amazon. That is, imagine they used toolbar/IE logs to infer that people went to amazon, searched for LCD TVs and then purchased model X. Then they could boost pages about X in search, or implement a "bestselling" feature. After all, they are "just" using clickstream data here. Similarly, they could track Netflix, and so on. IMHO, there's a loophole in Bing's a…

Absolutely on all of those sites. And Wikipedia. Why would I not want them to? The only reason I could think of that I would not want them to is if I thought it would create worse search links as a result. Although being able to personallize the search queries would be incredible. So when I search for "Movie XYZ" -- it can also look at my clickstream and see that I spend a lot of time in Netflix and Netflix has that…

I'm not talking about using just visits to improve search though, I'm talking about using the pattern of interaction on a site to basically replicate a site's data. That is, imagine netflix recommends movie X to you, now Bing infers that Netflix has done so using custom code to parse their IE logs, and then recommend movie X to you.

IMHO, this would be completely unethical.

That's exactly what Bing has admitted to doing with google: http://online.wsj.com/article/SB1000142405274870412450457611...

"Stefan Weitz, director of the Bing search engine at Microsoft, said in an interview the company studies how certain users interact with Google in order to improve Bing. ..."

Re: Matt Cutt's thoughts on Google Bing Debate

#34
post #7

Matt's assertion about "Suggested Sites" sending this data seems to be conjecture. I ran some packet captures and didn't find anything of that kind. However, if you install the Bing Toolbar then it does send URL clickstream data. It explicitly asks you beforehand if you want to send info about "the searches you do, websites you visit..." though. Full post: http://projectgus.com/2011/02/bing-google-finding-some-facts.…

Excellent analysis, thank you, especially in the distinction between "suggested sites" and the Bing toolbar's behavior. I can't argue with your methods. However, I do differ with some of your conclusions:

> The behaviour I’ve seen explains Google’s experiments, but does not support the accusation that Bing set out to copy Google.

I don't think it's so much about "set out to copy Google" necessarily, as it is that they are explicitly parsing Google queries and results from the clickthrough data and using it (quite directly) for their own results. What they set out to do is immaterial given what is provably happening.

> Bing Toolbar is tracking user clicks and Bing could use the result to improve search results. I don’t personally see any great distinction between this behaviour and Google’s many tracking, indexing and scraping endeavours which they use to improve their own search results.

The difference is that Google has proven that the results of certain queries are being directly fed as Bing results. If Microsoft does the same with Google rankings, I'd see your point, but right now the evidence only points in one direction.

> While I personally dislike the privacy implications, Bing Toolbar is pretty upfront about it when it gets installed (unlike much web page user tracking.) The fact that the tracking is plain HTTP not HTTPS, with the content in plaintext, would seem to indicate that they weren’t seeking to hide anything.

I'd be interested to see if Google over HTTPS queries are being transmitted by the toolbar over HTTP. That would be a pretty serious privacy violation IMO, especially when you pair that with unencrypted wifi at Starbucks. See the AOL search log fiasco: http://www.somethingawful.com/d/weekend-web/aol-search-log.p...

Re: Matt Cutt's thoughts on Google Bing Debate

#35
post #34
post #7

Matt's assertion about "Suggested Sites" sending this data seems to be conjecture. I ran some packet captures and didn't find anything of that kind. However, if you install the Bing Toolbar then it does send URL clickstream data. It explicitly asks you beforehand if you want to send info about "the searches you do, websites you visit..." though. Full post: http://projectgus.com/2011/02/bing-google-finding-some-facts.…

Excellent analysis, thank you, especially in the distinction between "suggested sites" and the Bing toolbar's behavior. I can't argue with your methods. However, I do differ with some of your conclusions: > The behaviour I’ve seen explains Google’s experiments, but does not support the accusation that Bing set out to copy Google. I don't think it's so much about "set out to copy Google" necessarily, as it is that the…

The difference is that Google has proven that the results of certain queries are being directly fed as Bing results. If Microsoft does the same with Google rankings, I'd see your point, but right now the evidence only points in one direction.

However, these are also the only experiments that have been run - outlier data, gaming the algorithms. If you only test one possible outlier scenario and don't control against any other, it's fairly ambitious to stand up and say "this is exactly what is happening!"

IMHO it would have been better for Microsoft to respond along these lines, instead of going into counterspin mode, though.

Re: Matt Cutt's thoughts on Google Bing Debate

#36

Earlier quoted context omitted.

Absolutely on all of those sites. And Wikipedia. Why would I not want them to? The only reason I could think of that I would not want them to is if I thought it would create worse search links as a result. Although being able to personallize the search queries would be incredible. So when I search for "Movie XYZ" -- it can also look at my clickstream and see that I spend a lot of time in Netflix and Netflix has that…

I'm not talking about using just visits to improve search though, I'm talking about using the pattern of interaction on a site to basically replicate a site's data. That is, imagine netflix recommends movie X to you, now Bing infers that Netflix has done so using custom code to parse their IE logs, and then recommend movie X to you. IMHO, this would be completely unethical. That's exactly what Bing has admitted to do…

I'm not sure I get your example (not clear when Bing would recommend a movie to me). So let me give that I do understand.

I go to Netflix and search for the move "Network". I end up clicking on "The Social Network". Later on Bing, if I search for "movie Network" -- I'd hope that one of the movies that comes back is "The Social Network", based on that clickthrough data from Netflix..

In my mind the only thing that is borderline unethical is that we've artificially limited this clickstream data to one company. I'd like to give it out more broadly. It's my data right?

Re: Matt Cutt's thoughts on Google Bing Debate

#37
post #7

Matt's assertion about "Suggested Sites" sending this data seems to be conjecture. I ran some packet captures and didn't find anything of that kind. However, if you install the Bing Toolbar then it does send URL clickstream data. It explicitly asks you beforehand if you want to send info about "the searches you do, websites you visit..." though. Full post: http://projectgus.com/2011/02/bing-google-finding-some-facts.…

Thank you. And this is precisely the problem with this whole fiasco--its just a bunch of conjecture and assertions. But because its from Google its taken with much more weight than it would be otherwise.

Google has created precisely the conversation they wanted based on their assertions then gets indignant when Microsoft refuses to step into it.

Re: Matt Cutt's thoughts on Google Bing Debate

#38
post #5

I think Matt lays out a strong argument for why they have made such a stink. And it sounds like a pretty compelling defense for their actions to go public. But... 1) I think they will regret it bigtime when all the attention they are causing makes the bored government officials poke their head in and realize that at a macro level, neither Bing nor Google does anything to protect user search privacy. 2) I think Google…

> 1) ...neither Bing nor Google does anything to protect user search privacy.

do you mean external or internal privacy? in terms of leaking information, SSL for search is about as good as you're going to get for privacy...if it's the browser or an extension (toolbar) that's watching searches, there's nothing a web site can do.

> 2) I think Google has more to lose by bringing this to light than they have to win.

That may be true, but there was a pretty cutting colbert segment on this last night where nuances about clickstreams weren't really a concern. personally I think google should have taken the humour route in the first place, as the "smoking gun" isn't all that damning at first sight. it requires some thought, which leaves plenty of room for disagreement and doubt.

Re: Matt Cutt's thoughts on Google Bing Debate

#39
post #5

I think Matt lays out a strong argument for why they have made such a stink. And it sounds like a pretty compelling defense for their actions to go public. But... 1) I think they will regret it bigtime when all the attention they are causing makes the bored government officials poke their head in and realize that at a macro level, neither Bing nor Google does anything to protect user search privacy. 2) I think Google…

I'm certain they never stopped trying to be innovative, but Bing is able to use that click stream data to reap the benefits of Google's innovations. Google is on the right side of this argument even if it makes them seem whiny to some people.

The part that seems interesting to me is that Google scrapes other entities' data (Reader, News, Books, Scholar) in dozens of other ways, and that behaviour is generally regarded as totally justifiable.

Re: Matt Cutt's thoughts on Google Bing Debate

#40

Matt Cutts should know better. Being part of '1000 signals' does not mean all signals are weighted evenly. It does not even mean the signals are weighted the same across all query types. This is machine learning - the actual weighting is learned and dynamic (always shifting) and not controlled. And there is absolutely no reason for Microsoft to take out a particular signal just because Google asked. There needs to be…

I don't think you read the whole article, because Matt quotes Nate Silver in saying that exact thing

You said: " -Being part of '1000 signals' does not mean all signals are weighted evenly."

Matt quotes Nate: "First, not all of the inputs are necessarily equal. It could be, for instance, that the Google results are weighted so heavily that they are as important as the other 999 inputs combined."

Post reply on HN