Live data from Hacker News

Matt Cutt's thoughts on Google Bing Debate

mattcutts.com

111–120 of 181 posts

Re: Matt Cutt's thoughts on Google Bing Debate

#112

Earlier quoted context omitted.

I'm not sure I get your example (not clear when Bing would recommend a movie to me). So let me give that I do understand. I go to Netflix and search for the move "Network". I end up clicking on "The Social Network". Later on Bing, if I search for "movie Network" -- I'd hope that one of the movies that comes back is "The Social Network", based on that clickthrough data from Netflix.. In my mind the only thing that is…

I was considering the case when you go to NetFlix, and they recomment the movie "Social Network" to you, and you click on it. Surely the url will have sufficient info to claim that its a recommendation. Now bing implements a movie recommendation service, which also magically recommends for you "Social Network", based on clicktracking of Netflix. IMHO, what Bing is doing with Google is exactly the same scenario as abo…

OK, you're talking about a new service. Gotcha.

Why wouldn't I want Bing to recommend the movie? The first thing Bing should do is use Netflix's APIs to suck down all of my ratings of movies. Then it should use my clickstream data for Netflix, IMDB, Hulu, Amazon, Facebook etc... And if I'm using Media Center or SageTV, I want those things to also get consolidated together.

Are you saying you want your Netflix data in one silo, your Hulu data elsewhere, your TV data elsewhere? Don't get me wrong, I want this to all be opt-in, but when I opt-in on MY behavior, I want them to use it.

Re: Matt Cutt's thoughts on Google Bing Debate

#113

Earlier quoted context omitted.

"Google created an artificial scenario where the ONLY input was Google search behavior and lo, the search results are exactly the same." Cutts addresses this: "As we said in our blog post, the whole reason we ran this test was because we thought this practice was happening for lots and lots of different queries, not simply rare queries."

If Google think they have other evidence, let's see it. The burden of proof is on the accuser, I think. (c/p from other thread.)

The problem is that the results would not be so dramatic. It would require statistical analysis to show that there was a difference. Most people's eyes would glaze over -- even though copying was occurring and genuine damage was being done (to market share, to business model, etc.), similar to a low-grade toxin in the environment.

Re: Matt Cutt's thoughts on Google Bing Debate

#114

Earlier quoted context omitted.

I don't understand how this isn't a settled issue. If clickstream data is 1 of 1000 signals, and you create clickstream data for a specialized query that will never trigger off another signal, then your created data will be reflected. That sounds exactly like what happened. You'd have to make the argument that using this data is wrong, somehow. But to make that argument, you'd basically have to argue that users shoul…

Accept in the "original" [torsoraphy] search. Forget Bing having to compete with Google's "spell correction team;" Bing need only use Google [copy] as a high frequency signal on tailing queries. That sounds like exactly what happened, and it's wrong.

Why exactly is it wrong?

Re: Matt Cutt's thoughts on Google Bing Debate

#115

My two favorite takeaways from this blog post are: 1) You can increase your site's rankings by installing the Bing Toolbar amongst a group of people and have said group search google for your target keyword and click through on your result. 2) Google is able and willing to manually screw around with search results in a seemingly easy manner. As to the issue of Bing slurping up google's data for their own purposes: Sc…

You can increase your site's rankings by installing the Bing Toolbar amongst a group of people and have said group search google for your target keyword and click through on your result. But you can just skip the toolbar altogether. Just go to Bing and do the search and click your site. Likewise, if you go to Google and do the same, you'll increase your ranking on Google. There is nothing new here, except if you inst…

Unless Bing weighs clicks on Google SERPs more strongly than clicks on Bing SERPs! After all, Google is the market leader...

This is the problem, we really don't know how hollow Bing is any more. If it's a thin UI over results that grow closer to Google, it's intellectually shallow. I wouldn't trust Bing search to combat, e.g., click fraud and SEO better than Google. If anything I would expect the edge cases and tricky behavior to derange Bing further.

Re: Matt Cutt's thoughts on Google Bing Debate

#116
post #96

Earlier quoted context omitted.

Web content layout and text owned by website. Opt out with robots.txt. Clickstream owned by user. Opt out by turning off toolbar. Both should be respected.

"Clickstream owned by user. Opt out by turning off toolbar." It is not clear to me that the user completely owns the clickstream. If Google's terms of use say you cannot share their search results with anyone else, are you legally allowed to send the clickstream to Microsoft?

I guess if it really comes down to it Google can encrypt the query in the URL. If Microsoft circumvents this (like by intercepting the JavaScript that encrypts it, or pulling the query out of the HTML input element), then it will demonstrate that they're really intentionally copying Google's results and this "clickstream" business was just a convenient cover.

Re: Matt Cutt's thoughts on Google Bing Debate

#117

Earlier quoted context omitted.

What is being data mined is a bit more than a user broadcasting their own preference on the correct result. The user is broadcasting a URL which is selected based on two factors: - the user's preference - Google's ranking algorithm. Had Google not ranked that URL, the user wouldn't be broadcasting it. If there was some way to extract the factor of the user's preference of URLs as a signal without the factor of Google…

Yup this is the issue Google wants to raise. Expect Google lobbying soon for some law along the lines of "clicks on somesite.com belong to somesite.com and can only be shared with another party with somesite.com's permission.". Then in some Google ToS the "user" will be granted access to the clicks that the user makes on Google's search results page... but third party apps like Bing toolbar won't have access.

A far-fetched hypothetical scenario that has little basis in fact.

Re: Matt Cutt's thoughts on Google Bing Debate

#118
post #5

I think Matt lays out a strong argument for why they have made such a stink. And it sounds like a pretty compelling defense for their actions to go public. But... 1) I think they will regret it bigtime when all the attention they are causing makes the bored government officials poke their head in and realize that at a macro level, neither Bing nor Google does anything to protect user search privacy. 2) I think Google…

Google has something to gain from all this. Consider the implications to Bing/Google competition for marketshare in the following scenario:

MS decides to bundle and auto-install the Bing toolbar in IE. Also default opt-in to the "share clickstream data with MS" option. Now you have tons of users with the Bing toolbar _and_ using Google to search. MS could "use" or "copy" the results of Google's hardwork on search and pagerank and consequently provide some serious competition for search marketshare.

Re: Matt Cutt's thoughts on Google Bing Debate

#119
post #87

Earlier quoted context omitted.

robots.txt

robots.txt is a good example of something that was introduced to make the new search engine technology more ethical but I'm not sure what point you are trying to make. If you're saying that we need a new robots.txt option saying "don't use clickstream data to help people find this page", I don't disagree; I'm not sure how many sites would take advantage of it though.

The point is that the creators of the data were given a say in how it was used. Googlers have busted a gut to provide 99% of the value of the clickstream data so they should determine how that data is used.

There's a more general point here: If something is difficult to do but easy to copy, then society should prevent that copying so that the creator can be rewarded for their effort. This maximises creativity and productivity and boosts GDP. Musicians, inventors, artists, authors, drug companies and software developers all rely on this principle. If we didn't have copyright and patents and robots.txt then there would be no incentive to produce half of the things people want, and we would all be much worse off. Google (and any other search engine) should be subject to the same rules.

Re: Matt Cutt's thoughts on Google Bing Debate

#120

Earlier quoted context omitted.

Then after a few weeks of sniffing clicks, Bing comes up with the same set of revolutionary results, but they have no idea how they got there, they have no idea what the evidence is to rank them there, all they know is that people like those results on Google. But this whole "no idea how they got there" is kind of bogus. The whole notion of link analysis is to infer from a link that a page is important. Now they're i…

> Its because the second order effect of a click implies relevance, just as a link does. There are two pieces of information in a click, one provided by google, and the other provided by the user. Google says, "'foo.com' is a good result for [foo]." The user says, "yup." Which is providing more information here? Everyone at google is happy to admit that if the internet didn't exist, we would have a very hard time ran…

Your first statement is wrong.

The user provides at least two pieces of information.

1) The search query (which you attmpted to hide in Google's information)

2) A URL they retrieve from Google

Google provides one piece:

1) An ordered list

By FAR the most important piece of information is the search query.

In fact, given a choice between having the search query or an arbitrary ordered list, I'm sure Bing would value the search query the most.

Clickstream data from Google is almost certainly but one piece of clickstream data (even if you are special-cased). Just as Microsoft.com is just one page you index. Now you wouldn't say that you'd have time ranking the internet if Microsoft.com didn't exist would you? Of course not.

And I think this is where Google is going astray. It's fine to call a spade a spade, but your overreaching. Trying to argue that MS couldn't index w/o Google is simply absurd. And if Matt was trying to be diplomatic, you're certainly not being by implying that MS would have a hard time indexing w/o Google.

Post reply on HN