Live data from Hacker News

Matt Cutt's thoughts on Google Bing Debate

mattcutts.com

71–80 of 181 posts

Re: Matt Cutt's thoughts on Google Bing Debate

#71

Earlier quoted context omitted.

I'm not talking about using just visits to improve search though, I'm talking about using the pattern of interaction on a site to basically replicate a site's data. That is, imagine netflix recommends movie X to you, now Bing infers that Netflix has done so using custom code to parse their IE logs, and then recommend movie X to you. IMHO, this would be completely unethical. That's exactly what Bing has admitted to do…

I'm not sure I get your example (not clear when Bing would recommend a movie to me). So let me give that I do understand. I go to Netflix and search for the move "Network". I end up clicking on "The Social Network". Later on Bing, if I search for "movie Network" -- I'd hope that one of the movies that comes back is "The Social Network", based on that clickthrough data from Netflix.. In my mind the only thing that is…

I was considering the case when you go to NetFlix, and they recomment the movie "Social Network" to you, and you click on it. Surely the url will have sufficient info to claim that its a recommendation.

Now bing implements a movie recommendation service, which also magically recommends for you "Social Network", based on clicktracking of Netflix.

IMHO, what Bing is doing with Google is exactly the same scenario as above (extract click data from a competitor's service, and surface the exact same data in their own product).

Re: Matt Cutt's thoughts on Google Bing Debate

#72

"Copying" was a pretty brutal word to use-- not surprising that it raised MSFT's hackles a bit. MS clearly uses toolbar users' clickstreams (on and off Google) to improve their own search efforts. Google created an artificial scenario where the ONLY input was Google search behavior and lo, the search results are exactly the same. Whether or not that steps over a line (I don't feel that it does), it's not "copying" in…

"Google created an artificial scenario where the ONLY input was Google search behavior and lo, the search results are exactly the same." Cutts addresses this: "As we said in our blog post, the whole reason we ran this test was because we thought this practice was happening for lots and lots of different queries, not simply rare queries."

If Google think they have other evidence, let's see it. The burden of proof is on the accuser, I think.

(c/p from other thread.)

Re: Matt Cutt's thoughts on Google Bing Debate

#73
The technical aspect is plain. Copying Google is a great idea. If you're tying to predict the weather and are given a general bag of indicators, the one that is a well-reasoned expert opinion of the weather is probably a very important source. It may even account for a majority of your own predictive power. It would be stupid to ignore it if your goal is to simply make the best prediction.

Likewise, aggressively scraping Google is a smart move. Then you add some new innovation atop it and have a real opportunity to return more informed responses. This is done all the time in science.

In some sense Google simply has to acknowledge that they are a pretty important segment of the web, not some separate entity from it.

So only the legal/ethical question remains. In science it's unethical to work atop someone else's project without crediting them. I doubt Bing would be interested in adding a Powered by Google bar. Moreover, since Bing could directly profit off Google's work undercutting actual algorithmic progress through pure marketing competition (hypothetically, anyway, I am sure that Bing has added tech too) I feel like it's better to restrict this sort of thing.

I think it's fair to say that much like commercial images on Flickr or sample songs, it is unethical and illegal to copy digital services and goods then either claim them your own or profit off of them. I think Google results are suitably close in spirit to this.

So maybe Bing and Girl Talk need to team up and discover and defend the ethical rights of sampling digital goods.

Re: Matt Cutt's thoughts on Google Bing Debate

#74
post #5

I think Matt lays out a strong argument for why they have made such a stink. And it sounds like a pretty compelling defense for their actions to go public. But... 1) I think they will regret it bigtime when all the attention they are causing makes the bored government officials poke their head in and realize that at a macro level, neither Bing nor Google does anything to protect user search privacy. 2) I think Google…

> They should focus on trying to be innovative again, that was the Google I respected. I work in search quality at Google. I'm busting my ass every day working on fundamental reimaginings of how results get ranked. I'm going to keep doing that regardless of what bing does, because it makes the world a better place, and it's fun. But, suppose the stuff I'm working on works out, and tomorrow Google shows up with a whol…

Is this fair? Is this ethical? Is this even legal?

Oh please. At worst, Microsoft's customers are sharing results of their browsing activities, based on information you gave them. How in the world could that be illegal?

Your team is really coming across as whiny on this issue, which seems to me precisely the wrong tactic to take in response to Bing's increasing market share. It makes your team look defensive, as if your advantage is slipping away and this sort of theatrical display is all you've got in response.

Re: Matt Cutt's thoughts on Google Bing Debate

#75
post #5

I think Matt lays out a strong argument for why they have made such a stink. And it sounds like a pretty compelling defense for their actions to go public. But... 1) I think they will regret it bigtime when all the attention they are causing makes the bored government officials poke their head in and realize that at a macro level, neither Bing nor Google does anything to protect user search privacy. 2) I think Google…

> They should focus on trying to be innovative again, that was the Google I respected. I work in search quality at Google. I'm busting my ass every day working on fundamental reimaginings of how results get ranked. I'm going to keep doing that regardless of what bing does, because it makes the world a better place, and it's fun. But, suppose the stuff I'm working on works out, and tomorrow Google shows up with a whol…

Keep doing what you are doing and don't take what I wrote too personally. It's obvious that there are a lot of amazingly talented individuals working hard at Google.

IMHO - One of your colleagues should have stopped that blog post from being put up.

Re: Matt Cutt's thoughts on Google Bing Debate

#76
Seems like Bing is having trouble explaining exactly what they are doing, and Google is having trouble explaining why they should stop.

Is it illegal? If not, why should Bing stop doing it?

> I think Bing’s engineers deserve to know that when they beat Google on a query, it’s due entirely to their hard work. Unless Microsoft changes its practices, there will always be a question mark.

Kind of rings hollow to me. If I were Bing, I'd want to do what's best for my users.

Re: Matt Cutt's thoughts on Google Bing Debate

#77
post #5

I think Matt lays out a strong argument for why they have made such a stink. And it sounds like a pretty compelling defense for their actions to go public. But... 1) I think they will regret it bigtime when all the attention they are causing makes the bored government officials poke their head in and realize that at a macro level, neither Bing nor Google does anything to protect user search privacy. 2) I think Google…

> They should focus on trying to be innovative again, that was the Google I respected. I work in search quality at Google. I'm busting my ass every day working on fundamental reimaginings of how results get ranked. I'm going to keep doing that regardless of what bing does, because it makes the world a better place, and it's fun. But, suppose the stuff I'm working on works out, and tomorrow Google shows up with a whol…

Then after a few weeks of sniffing clicks, Bing comes up with the same set of revolutionary results, but they have no idea how they got there, they have no idea what the evidence is to rank them there, all they know is that people like those results on Google.

But this whole "no idea how they got there" is kind of bogus. The whole notion of link analysis is to infer from a link that a page is important. Now they're infering from a click that page is important. The relative importance of that page is a function of the search query, like the relative importance of a link is a function of the page from which it originated.

Link analysis and click throughs are both second order effects, right? If CNN.com starts linking to a page, you don't really know why that page is important, except for the fact that CNN.com now points to it.

So lets be clear MS knows how it got there as much as link analysis does. Its because the second order effect of a click implies relevance, just as a link does.

Re: Matt Cutt's thoughts on Google Bing Debate

#78
post #67

Earlier quoted context omitted.

That's not how the AFP felt about Google News when it launched: http://en.wikipedia.org/wiki/Google_News#News_agencies or how some publishers felt about Google Books when it launched: http://en.wikipedia.org/wiki/Google_Books#Copyright_infringe...

The AFP could have opted out if they wanted to. Copyright and fair use laws are not something we should be dragging into this. Your examples have nothing to do with the argument at hand between Google and Bing.

I'd suggest that they're indirectly relevant, because they all address the question of what constitutes fair indexing and what constitutes "copying" or theft of innovative/valuable information.

The evidence suggests Bing is scraping clickstream data to work out where its users go and where they stay, and this includes google search URLs with embedded queries. The evidence doesn't seem to suggest that Bing treats google any differently to any other URL on the web.

So, the questions could come down to: who has the right to disseminate URL strings and mine them for valuable information? Which seems not indirectly related to "who has the right to digitise other forms of media and mine them for valuable information?"

Re: Matt Cutt's thoughts on Google Bing Debate

#79
post #50

Earlier quoted context omitted.

If this joint toolbar comes to fruition, which it never will, the crowd will decide the obvious winner for practically all queries. At that point what will be the difference between Bing and Google?

At that point what will be the difference between Bing and Google? Speed of innovation. It's not our job to make it easy for them to have diffentiation. I just want great results. The more parity we give them with respect to data, the harder they'll have work to innovate in order to differentiate. AFAICT, that benefits us the users.

I didn't suggest it should be our job to make it easy for them. I want great results too and I want it without ads and spam getting in the way as well. In fact there should be more tools integrated into browsers to allow the user to own their own click stream and browsing pattern data and "sell" it to the highest bidder. Do you know of such tools? I don't. We are even at the point both in terms of computational power and data storage costs that each user could potentially have their own web crawler/search engine. The bing and google bars basically steal this information from users and in return provide little to no improvement on search results. The point I was making was that there is absolutely no incentive on either side to share knowledge and converge onto a high quality, common query result set because both companies try really really hard to not appear like the other guy and the result set is part of this.

Re: Matt Cutt's thoughts on Google Bing Debate

#80
A quick thought experiment on this subject: say I search for a term on Bing, find a link I want and then put that link on my blog - when Google indexes that blog have they copied Bing?

Sure its different, but is it meaningfully different? I made the link between the two terms, I also consented for that data to be used in both cases (assuming the data comes from the Bing toolbar and an agreeable robots.txt). I just don't see how the data that Microsoft is using is off limits.

Post reply on HN