Live data from Hacker News

Matt Cutt's thoughts on Google Bing Debate

mattcutts.com

151–160 of 181 posts

Re: Matt Cutt's thoughts on Google Bing Debate

#151
post #149

Earlier quoted context omitted.

> They should focus on trying to be innovative again, that was the Google I respected. I work in search quality at Google. I'm busting my ass every day working on fundamental reimaginings of how results get ranked. I'm going to keep doing that regardless of what bing does, because it makes the world a better place, and it's fun. But, suppose the stuff I'm working on works out, and tomorrow Google shows up with a whol…

You have computed a relationship between FOO and BAR. For as long as you keep it to yourself, this ephemeral relationship is a secret, yours to keep. Once you tell the world that FOO and BAR are related, and the whole world looks at it and says "Yay, it IS related!", the relationship stops being ephemeral and becomes actual, reflected in actions of the users. You no loner own that relationship. You can't claim owners…

>You can't claim ownership of facts, even if you discovered them first, or created them into existence.

It's not as straightforward as all that. Suppose you write a book. The existence of that book is a fact. The words that are in that book are now a fact. Now, I create a book that says, "The following book was written by DennisM, this is a list of the words that are in that book, in order" and then precede to duplicate your book. Those are the words that are in the book, it's a fact, no one can deny that it's a fact. That's not defensible, even though I'm just stating a fact.

We're in a somewhat grey area here, dealing with ethical issues that have never been clearly dealt with before. However, I feel that a general principle of the web applies: If I create some data, and don't at least implicitly give you permission to use it, it's unethical for you to use it.

This is the central idea behind copyright, behind robots.txt, behind plagiarism, and even behind privacy.

Re: Matt Cutt's thoughts on Google Bing Debate

#152

Earlier quoted context omitted.

The reason is that with the URL and the query, there's a decent chance you can derive the association. Now the reason I say this is that to my best approximation the key difference that makes Google better than Bing (when it is) is the index size, and not the search algorithm. So the two key pieces of value are: A) queries that people don't do on your site, which would have generated poor relevance B) Pages that are…

>is the index size, and not the search algorithm I'm pretty thoroughly sure this is false. I'm sorry I can't show data to back it up, but Google has many many systems that are years ahead of bing's technology. I'm kinda hamstrung here by being unable to reveal anything about them. To a decent approximation, both Google and Bing likely have the whole internet that matters (and that they are allowed to crawl) in their…

One obvious and very public example of how the algorithm matters more than the index size would be Cuil. It was launched with lots of hype about how it would have a index several times larger than Google's. And of course the results were notoriously bad.

I'm stunned that anyone could think that the core issue in this controversy is indexing or page discovery. It should have been totally obvious from the examples in the initial article and blog post that this was specifically about copying ranking.

Re: Matt Cutt's thoughts on Google Bing Debate

#153
post #134

Earlier quoted context omitted.

If Google think they have other evidence, let's see it. The burden of proof is on the accuser, I think. (c/p from other thread.)

If this isn't enough to convince you, nothing will be. I'm not sure how the evidence could be any more convincing. Content from Google's result database ends up used to generate Bing's results, full stop. Proved. It may be technical error, it may be only used a little bit, it may be deliberate, it may be accidental and perfectly ethical by some people's standards, there's all kinds of ways to legitimately interpret t…

"... we thought this practice was happening for lots and lots of different queries, not simply rare queries ..."

My comment was in reference to that.

Re: Matt Cutt's thoughts on Google Bing Debate

#154

Earlier quoted context omitted.

Imagine I launch a search engine with no data. Then I feed it with urls IE users click on after their google search. I will eventually end up with an exact copy of google database. So I think that this technique can be called "copying". Now if Bing uses this technique for 0.1% of their data, then it can be said that 0.1% of their data are copied from Google database.

What do you suppose google does when they are faced with a novel query and their algorithm returns 10 equally good results as matches? At that point you might as well provide a whole bunch of users some permutation of the matches and take into account which links are clicked on the most. This is exactly what bing is doing and if you crawl the web then you will indeed end up with google's database, the only difference…

That's not what bing is doing in this case. What bing is doing is taking information from Google, as collected by users, and presenting that information to their own users.

Re: Matt Cutt's thoughts on Google Bing Debate

#155

Earlier quoted context omitted.

Imagine I launch a search engine with no data. Then I feed it with urls IE users click on after their google search. I will eventually end up with an exact copy of google database. So I think that this technique can be called "copying". Now if Bing uses this technique for 0.1% of their data, then it can be said that 0.1% of their data are copied from Google database.

That's a straw man argument. That's not what they did. They had a search engine with data (lots of it), using many of the same factors that Google does (page/domain authority, on-page markup, etc). The feed that algorithm a lot of data. One bit of data they feed it is search behvior of toolbar users (presumably ALL search engines). In your scenario, that'd be a copy. In the reality scenario (described above), I think…

search behvior of toolbar users (presumably ALL search engines)

The situation is even less compelling than this. There's no evidence to suggest it's not the complete navigation behaviour of toolbar users, not just when they're on search engines.

Re: Matt Cutt's thoughts on Google Bing Debate

#156
post #90

Earlier quoted context omitted.

> They should focus on trying to be innovative again, that was the Google I respected. I work in search quality at Google. I'm busting my ass every day working on fundamental reimaginings of how results get ranked. I'm going to keep doing that regardless of what bing does, because it makes the world a better place, and it's fun. But, suppose the stuff I'm working on works out, and tomorrow Google shows up with a whol…

You have a valid point that the click-through measures more of what Google puts out rather than what users find useful. But Bing has access not just to the click-through rate but also the rest of the clickstream. IF (and that's an all-caps if) the value of the click is weighted by engagement metrics on the resulting page, the association is more closely related to the value the user finds in the page rather than the…

[deleted]

Re: Matt Cutt's thoughts on Google Bing Debate

#157
post #87

Earlier quoted context omitted.

robots.txt is a good example of something that was introduced to make the new search engine technology more ethical but I'm not sure what point you are trying to make. If you're saying that we need a new robots.txt option saying "don't use clickstream data to help people find this page", I don't disagree; I'm not sure how many sites would take advantage of it though.

You're addressing the letter of the "law" not the spirit. Do we really need a new formal standard to indicate that this is unethical? To me it's pretty plain from the standards the web has already agreed upon.

Absolutely a new standard is needed. This is the first time any major player has suggested that data from a users clickstream should be covered by robots.txt or anything like it.

The comparison to a web directory is perfect - in the absence of a robots.txt Google and other crawlers think the data is absolutely fair game. There isn't yet a corresponding analogy for clickstream data so its fair game. Are you sure all of Google's tools respect robots.txt for passive analysis of user initiated actions on other websites?

Re: Matt Cutt's thoughts on Google Bing Debate

#158

My question is... did Google install the Bing toolbar and then performing these searches on Google and click the links? If not, is Bing suggesting that users searched for those convoluted strings, in IE, with the Bing toolbar, on Google's website, and then clicked those links? I hope this doesn't sound hyperbolic, that is really my understanding of the issue. Sorry, I'm not a big SEO guy, it's very possible my unders…

Really, -1? When I asked a question and prefaced it with a huge disclaimer stating that I might not understand the full implications of how Bing is using their toolbar metrics?

I just don't understand how Bing can claim they were using toolbar metrics when it seems highly unlikely that any user would be searching for these keywords. Let alone on Google, when they have the Bing toolbar installed.

But really, thanks for the anonymous downvotes.

Re: Matt Cutt's thoughts on Google Bing Debate

#159

A quick thought experiment on this subject: say I search for a term on Bing, find a link I want and then put that link on my blog - when Google indexes that blog have they copied Bing? Sure its different, but is it meaningfully different? I made the link between the two terms, I also consented for that data to be used in both cases (assuming the data comes from the Bing toolbar and an agreeable robots.txt). I just do…

Just putting the link on your blog doesn't establish the connection between the keywords and the link. If you put both the keywords and the link on your blog then it allows Google to independently establish that connection using the merits of their own algorithms.

The "incrimminating" part here is that Microsoft appears to be intentionally parsing the keywords out of the search. Which means they are intentionally looking at a click and saying a) this is a google search and here are the terms used and then b) this is a result that Google returned, and they are then using that to fill their own index. If they generically parsed the URL for terms then you might argue they are not giving Google special treatment, they are just doing this for every page the user goes to. However that's a bit hard to buy - if they really did that they would end up with all kinds of garbage associations from opaque URLs. So they must have a signal saying "this was a google search, treat it better than the others". Or they are somewhere in between the two. It's not clear to me where they fall on this scale - it's generally murky. At worst, I'd say they are copying, at best, I'd say it's sneaky but clever and fair game. The minute they single out Google and say "hey, this must be a good result" I think they crossed a line.

Re: Matt Cutt's thoughts on Google Bing Debate

#160
post #149

Earlier quoted context omitted.

You have computed a relationship between FOO and BAR. For as long as you keep it to yourself, this ephemeral relationship is a secret, yours to keep. Once you tell the world that FOO and BAR are related, and the whole world looks at it and says "Yay, it IS related!", the relationship stops being ephemeral and becomes actual, reflected in actions of the users. You no loner own that relationship. You can't claim owners…

> You can't claim ownership of facts, even if you discovered them first, or created them into existence. It's not as straightforward as all that. Suppose you write a book. The existence of that book is a fact. The words that are in that book are now a fact. Now, I create a book that says, "The following book was written by DennisM, this is a list of the words that are in that book, in order" and then precede to dupli…

Actually, it is as straightforward as all that. What's alleged is that someone searches for DenisM's book on Google, and when they click on one of the results, just that piece of data is forwarded to Microsoft. Not a scraping of all of Google's search results and nothing like the contents of the book.

As to "the central idea of copyright", bear in mind that in addition to not being able to copyright ideas or data, there's also the notion of "fair use".

Post reply on HN