Live data from Hacker News

Matt Cutt's thoughts on Google Bing Debate

mattcutts.com

101–110 of 181 posts

Re: Matt Cutt's thoughts on Google Bing Debate

#101
post #96

Earlier quoted context omitted.

"Clickstream owned by user. Opt out by turning off toolbar." It is not clear to me that the user completely owns the clickstream. If Google's terms of use say you cannot share their search results with anyone else, are you legally allowed to send the clickstream to Microsoft?

If Google's terms of use say you cannot share their search results with anyone else, are you legally allowed to send the clickstream to Microsoft? So if my dad calls with a question, and I search for the answer, I can't send him the link that I got from Google? Or do you mean something else?

He means something else. Sending the link is not the same as sending the series of clicks and inputs required to get to that link in Google.

Now it may be the case that sending clickstream data on Google to Microsoft is perfectly fine, but it's not exactly clear to me.

Re: Matt Cutt's thoughts on Google Bing Debate

#102
post #86

Earlier quoted context omitted.

I don't understand how this isn't a settled issue. If clickstream data is 1 of 1000 signals, and you create clickstream data for a specialized query that will never trigger off another signal, then your created data will be reflected. That sounds exactly like what happened. You'd have to make the argument that using this data is wrong, somehow. But to make that argument, you'd basically have to argue that users shoul…

You should be able to share your results with whoever you want but that's not what happens. You share your data with just bing and google. You don't really own any of the data you are generating so saying you should be able to share it with whoever you want is saying hop before you've jumped.

Where I click on a screen, and what I click on, is not my data?

Re: Matt Cutt's thoughts on Google Bing Debate

#103

"Copying" was a pretty brutal word to use-- not surprising that it raised MSFT's hackles a bit. MS clearly uses toolbar users' clickstreams (on and off Google) to improve their own search efforts. Google created an artificial scenario where the ONLY input was Google search behavior and lo, the search results are exactly the same. Whether or not that steps over a line (I don't feel that it does), it's not "copying" in…

Imagine I launch a search engine with no data. Then I feed it with urls IE users click on after their google search. I will eventually end up with an exact copy of google database. So I think that this technique can be called "copying". Now if Bing uses this technique for 0.1% of their data, then it can be said that 0.1% of their data are copied from Google database.

What do you suppose google does when they are faced with a novel query and their algorithm returns 10 equally good results as matches? At that point you might as well provide a whole bunch of users some permutation of the matches and take into account which links are clicked on the most. This is exactly what bing is doing and if you crawl the web then you will indeed end up with google's database, the only difference will be how you rank results.

Re: Matt Cutt's thoughts on Google Bing Debate

#104

Let’s take that thought to its conclusion. If clicks on Google really account for only 1/1000th (or some other trivial fraction) of Microsoft’s relevancy, why not just stop using those clicks and reduce the negative coverage and perception of this? And if Microsoft is unwilling to stop incorporating Google’s clicks in Bing’s rankings, doesn’t that argue that Google’s clicks account for much more than 1/1000th of Bing…

What is being data mined is a bit more than a user broadcasting their own preference on the correct result. The user is broadcasting a URL which is selected based on two factors: - the user's preference - Google's ranking algorithm. Had Google not ranked that URL, the user wouldn't be broadcasting it. If there was some way to extract the factor of the user's preference of URLs as a signal without the factor of Google…

Yup this is the issue Google wants to raise. Expect Google lobbying soon for some law along the lines of "clicks on somesite.com belong to somesite.com and can only be shared with another party with somesite.com's permission.". Then in some Google ToS the "user" will be granted access to the clicks that the user makes on Google's search results page... but third party apps like Bing toolbar won't have access.

Re: Matt Cutt's thoughts on Google Bing Debate

#105

"Copying" was a pretty brutal word to use-- not surprising that it raised MSFT's hackles a bit. MS clearly uses toolbar users' clickstreams (on and off Google) to improve their own search efforts. Google created an artificial scenario where the ONLY input was Google search behavior and lo, the search results are exactly the same. Whether or not that steps over a line (I don't feel that it does), it's not "copying" in…

Imagine I launch a search engine with no data. Then I feed it with urls IE users click on after their google search. I will eventually end up with an exact copy of google database. So I think that this technique can be called "copying". Now if Bing uses this technique for 0.1% of their data, then it can be said that 0.1% of their data are copied from Google database.

That's a straw man argument. That's not what they did. They had a search engine with data (lots of it), using many of the same factors that Google does (page/domain authority, on-page markup, etc). The feed that algorithm a lot of data. One bit of data they feed it is search behvior of toolbar users (presumably ALL search engines).

In your scenario, that'd be a copy. In the reality scenario (described above), I think it's not.

Re: Matt Cutt's thoughts on Google Bing Debate

#106
post #86

Earlier quoted context omitted.

You should be able to share your results with whoever you want but that's not what happens. You share your data with just bing and google. You don't really own any of the data you are generating so saying you should be able to share it with whoever you want is saying hop before you've jumped.

Where I click on a screen, and what I click on, is not my data?

Of course not. Try inquiring google or bing about obtaining all the information they have gathered from your toolbar and see what response you get. Actually I'll tell you: "The gathered data is anonymous and we have no way of verifying who sent it to us". So how exactly do you own information that you have absolutely no access to and with no way to transfer to competing search companies?

Re: Matt Cutt's thoughts on Google Bing Debate

#107
post #87

Earlier quoted context omitted.

robots.txt

robots.txt is a good example of something that was introduced to make the new search engine technology more ethical but I'm not sure what point you are trying to make. If you're saying that we need a new robots.txt option saying "don't use clickstream data to help people find this page", I don't disagree; I'm not sure how many sites would take advantage of it though.

You're addressing the letter of the "law" not the spirit. Do we really need a new formal standard to indicate that this is unethical? To me it's pretty plain from the standards the web has already agreed upon.

Re: Matt Cutt's thoughts on Google Bing Debate

#108

"Copying" was a pretty brutal word to use-- not surprising that it raised MSFT's hackles a bit. MS clearly uses toolbar users' clickstreams (on and off Google) to improve their own search efforts. Google created an artificial scenario where the ONLY input was Google search behavior and lo, the search results are exactly the same. Whether or not that steps over a line (I don't feel that it does), it's not "copying" in…

I don't understand how this isn't a settled issue. If clickstream data is 1 of 1000 signals, and you create clickstream data for a specialized query that will never trigger off another signal, then your created data will be reflected. That sounds exactly like what happened. You'd have to make the argument that using this data is wrong, somehow. But to make that argument, you'd basically have to argue that users shoul…

Accept in the "original" [torsoraphy] search. Forget Bing having to compete with Google's "spell correction team;" Bing need only use Google [copy] as a high frequency signal on tailing queries.

That sounds like exactly what happened, and it's wrong.

Re: Matt Cutt's thoughts on Google Bing Debate

#110

Earlier quoted context omitted.

> They should focus on trying to be innovative again, that was the Google I respected. I work in search quality at Google. I'm busting my ass every day working on fundamental reimaginings of how results get ranked. I'm going to keep doing that regardless of what bing does, because it makes the world a better place, and it's fun. But, suppose the stuff I'm working on works out, and tomorrow Google shows up with a whol…

Then after a few weeks of sniffing clicks, Bing comes up with the same set of revolutionary results, but they have no idea how they got there, they have no idea what the evidence is to rank them there, all they know is that people like those results on Google. But this whole "no idea how they got there" is kind of bogus. The whole notion of link analysis is to infer from a link that a page is important. Now they're i…

>Its because the second order effect of a click implies relevance, just as a link does.

There are two pieces of information in a click, one provided by google, and the other provided by the user.

Google says, "'foo.com' is a good result for [foo]."

The user says, "yup."

Which is providing more information here?

Everyone at google is happy to admit that if the internet didn't exist, we would have a very hard time ranking it. Is bing ready to admit that if google didn't exist, they'd have a hard time ranking the internet?

Post reply on HN