Live data from Hacker News

Matt Cutt's thoughts on Google Bing Debate

mattcutts.com

171–180 of 181 posts

Re: Matt Cutt's thoughts on Google Bing Debate

#171
post #134

Earlier quoted context omitted.

If this isn't enough to convince you, nothing will be. I'm not sure how the evidence could be any more convincing. Content from Google's result database ends up used to generate Bing's results, full stop. Proved. It may be technical error, it may be only used a little bit, it may be deliberate, it may be accidental and perfectly ethical by some people's standards, there's all kinds of ways to legitimately interpret t…

"... we thought this practice was happening for lots and lots of different queries, not simply rare queries ..." My comment was in reference to that.

I know. My point is that they've established beyond a shadow of a doubt that Google information is ending up in Bing's results. The default presumption is that any data could end up anywhere; the odd argument is that they have specially cased this and they only grab clickstream data when they have no other sources. One would expect that if the data is being collected it's being fed into the search algorithms and always being considered as a weight, not just sometimes.

I hate to be a bit harsh, but that argument sounds more like spin or rationalization than a logical argument. It sounds nice but it doesn't make sense if you try to actually map it back to the real world, where somebody had to type real code to produce the effect you're hypothesizing, and they weren't writing this code with this debate in mind in advance.

Re: Matt Cutt's thoughts on Google Bing Debate

#172
post #6

Oh man. That research paper he quoted strongly indicates that MS is specifically targeting Google. I was totally wrong. Bing are copying Google. I'm sorry (I am the "What on earth are Google thinking" author). I still think the honeypot experiments didn't support the conclusion. But this paper coupled with Bing's lacklustre pseudo-denial strongly indicates that my view of events was not accurate and that Bing were in…

user24, I didn't make the connection that you did those "What on earth are Google thinking" posts--just wanted to say thanks for your tweet earlier. It made me feel like doing the post wasn't such a bad idea.

Re: Matt Cutt's thoughts on Google Bing Debate

#173

That shows a lot of bad faith on the part of Matt Cutts. It has been explained to death, here and elsewhere, what most likely happened. The fact that he continues to make the same accusation... well frankly he lost a lot of credibility with me. One more manipulative corporate drone, one less genuine hacker.

Sorry you felt that way; was there a part of the post that struck you as especially off-base?

Re: Matt Cutt's thoughts on Google Bing Debate

#174
post #6

Oh man. That research paper he quoted strongly indicates that MS is specifically targeting Google. I was totally wrong. Bing are copying Google. I'm sorry (I am the "What on earth are Google thinking" author). I still think the honeypot experiments didn't support the conclusion. But this paper coupled with Bing's lacklustre pseudo-denial strongly indicates that my view of events was not accurate and that Bing were in…

user24, I didn't make the connection that you did those "What on earth are Google thinking" posts--just wanted to say thanks for your tweet earlier. It made me feel like doing the post wasn't such a bad idea.

I just wish you'd found that research right in the beginning! It's still not 100% conclusive, but it adds a lot more weight to your claims.

But still, I don't understand why Bing felt it necessary to specifically target you! The generic approach of mining all URLs indiscriminately, which I advocated, would have had largely the same impact and would be defensible against claims of cheating. It's crazy. There was no need to copy directly from Google and yet it looks very much like they still did.

Re: Matt Cutt's thoughts on Google Bing Debate

#175
post #90

Earlier quoted context omitted.

> They should focus on trying to be innovative again, that was the Google I respected. I work in search quality at Google. I'm busting my ass every day working on fundamental reimaginings of how results get ranked. I'm going to keep doing that regardless of what bing does, because it makes the world a better place, and it's fun. But, suppose the stuff I'm working on works out, and tomorrow Google shows up with a whol…

You have a valid point that the click-through measures more of what Google puts out rather than what users find useful. But Bing has access not just to the click-through rate but also the rest of the clickstream. IF (and that's an all-caps if) the value of the click is weighted by engagement metrics on the resulting page, the association is more closely related to the value the user finds in the page rather than the…

They are free to do that on their own search results. If they can't figure out how to rank the document in the first place, they shouldn't be getting data about the association from google.

Re: Matt Cutt's thoughts on Google Bing Debate

#176
post #171

Earlier quoted context omitted.

"... we thought this practice was happening for lots and lots of different queries, not simply rare queries ..." My comment was in reference to that.

I know. My point is that they've established beyond a shadow of a doubt that Google information is ending up in Bing's results. The default presumption is that any data could end up anywhere; the odd argument is that they have specially cased this and they only grab clickstream data when they have no other sources. One would expect that if the data is being collected it's being fed into the search algorithms and alwa…

Of course it's always being considered, but the question is how much effect it has in practice. I guess this depends on your ontological categories, but personally I distinguish between "Bing is copying (if you want to call it that, and fair enough) certain Google results" and "Bing is a copy of Google" (which I think has been alleged or insinuated, but not shown). The difference is a matter of quantity turning into quality.

I think they probably have statistics that convince them of the latter, and they should show them, so we can see if people outside Google also find them convincing.

Re: Matt Cutt's thoughts on Google Bing Debate

#177
post #130

Earlier quoted context omitted.

Accept in the "original" [torsoraphy] search. Forget Bing having to compete with Google's "spell correction team;" Bing need only use Google [copy] as a high frequency signal on tailing queries. That sounds like exactly what happened, and it's wrong.

Look at it this way: Google is setting trends on tail terms due to its massive market share. Once the trend is set, Bing captures the trend by monitoring user's behavior and provides that to the other users. By reflecting user's behavior in its own index, Bing seems to replicate the Google's index structure which created that behavior. Had there been no Google, there would have been different trends and Bing would ca…

This comment really gets to the meat of the issue here. Search is about discovering and then predicting user trends. Google in fact is creating user trends through its position in the market. Should Bing be locked out of the market because Google is currently in the dominant position?

If Bing is allowed to discover user trends then it will necessarily end up being a replica of Google's results in some cases. There's just no way around it.

Re: Matt Cutt's thoughts on Google Bing Debate

#178
post #154

Earlier quoted context omitted.

That's not what bing is doing in this case. What bing is doing is taking information from Google, as collected by users, and presenting that information to their own users.

Ya, and? They are also taking information from a whole bunch of other sites collected by their users and incorporating that information into ranking search results. I don't really see how this is inherently wrong.

I'm not saying that it's inherently wrong, I was attempting to explain to you how your example (Google using A/B testing to determine which results are the most relevant) had no relation to what bing is doing.

Re: Matt Cutt's thoughts on Google Bing Debate

#179
post #144

Earlier quoted context omitted.

Do we know they haven't? We only know they haven't reported on such a test.

Why would they purposefully make their case inconclusive?

There are three reasons for Google not to report results of a control-case test:

1) It didn't occur to them to run a control in their experiment.

2) It did occur to them but they decided not to do it for some reason.

3) They did run the control but it did not bolster their case.

As you point out, if they ran it and it bolstered their case there would be no reason not to report it. Of the above reasons, #2 seems the least likely (because it's easy to set up a control and it would better clarify the situation). I do not have enough evidence to judge whether #1 or #3 is more likely.

Re: Matt Cutt's thoughts on Google Bing Debate

#180
post #149

Earlier quoted context omitted.

You have computed a relationship between FOO and BAR. For as long as you keep it to yourself, this ephemeral relationship is a secret, yours to keep. Once you tell the world that FOO and BAR are related, and the whole world looks at it and says "Yay, it IS related!", the relationship stops being ephemeral and becomes actual, reflected in actions of the users. You no loner own that relationship. You can't claim owners…

> You can't claim ownership of facts, even if you discovered them first, or created them into existence. It's not as straightforward as all that. Suppose you write a book. The existence of that book is a fact. The words that are in that book are now a fact. Now, I create a book that says, "The following book was written by DennisM, this is a list of the words that are in that book, in order" and then precede to dupli…

In your analogy there was a wholesale copyright infringement, entirely absent from bing sting. Here's a much closer analogy:

I wrote a book, in which I described some new way to dance. Before the book was published, the ideas were mine to keep secret. After the book was published, a bunch of readers picked up on the idea. They started all dancing in certain way, and someone described their behavior. That description will be de-facto copy of my book, and yet it is entirely legitimate. Because users own they behavior, not the author of the book who inspired them.

Now if someone simply copied pages from my book, that would be a copyright violation. But that's not what happened. As it is, it's an original work of art. Would that suck for me as a dance-inventor? Obviously. Do I want to live in a society where description of my behavior is owned by the person who inspired or directed it? Absolutely not.

Your idea of data ownership is contrary to tradition and contrary to the law. Facts (such as relationship between FOO and BAR) can not be owned in any way shape or form, except as a trade secret. You can not copyright a fact, or trademark a fact. In some cases you can patent application of a fact to a problem, but that's not the case here, as you can't and don't want to patent relationship between a word and target page.

Just because you put effort into something, does not mean it's yours. You probably wish it were true, but again, that's not how the law works, and not what the tradition is. You can't own facts.

Post reply on HN