"Copying" was a pretty brutal word to use-- not surprising that it raised MSFT's hackles a bit. MS clearly uses toolbar users' clickstreams (on and off Google) to improve their own search efforts. Google created an artificial scenario where the ONLY input was Google search behavior and lo, the search results are exactly the same. Whether or not that steps over a line (I don't feel that it does), it's not "copying" in…
I don't understand how this isn't a settled issue. If clickstream data is 1 of 1000 signals, and you create clickstream data for a specialized query that will never trigger off another signal, then your created data will be reflected. That sounds exactly like what happened. You'd have to make the argument that using this data is wrong, somehow. But to make that argument, you'd basically have to argue that users shoul…
Matt Cutt's thoughts on Google Bing Debate
161–170 of 181 posts
Re: Matt Cutt's thoughts on Google Bing Debate
#162Earlier quoted context omitted.
Imagine I launch a search engine with no data. Then I feed it with urls IE users click on after their google search. I will eventually end up with an exact copy of google database. So I think that this technique can be called "copying". Now if Bing uses this technique for 0.1% of their data, then it can be said that 0.1% of their data are copied from Google database.
That's a straw man argument. That's not what they did. They had a search engine with data (lots of it), using many of the same factors that Google does (page/domain authority, on-page markup, etc). The feed that algorithm a lot of data. One bit of data they feed it is search behvior of toolbar users (presumably ALL search engines). In your scenario, that'd be a copy. In the reality scenario (described above), I think…
Re: Matt Cutt's thoughts on Google Bing Debate
#163A quick thought experiment on this subject: say I search for a term on Bing, find a link I want and then put that link on my blog - when Google indexes that blog have they copied Bing? Sure its different, but is it meaningfully different? I made the link between the two terms, I also consented for that data to be used in both cases (assuming the data comes from the Bing toolbar and an agreeable robots.txt). I just do…
Just putting the link on your blog doesn't establish the connection between the keywords and the link. If you put both the keywords and the link on your blog then it allows Google to independently establish that connection using the merits of their own algorithms. The "incrimminating" part here is that Microsoft appears to be intentionally parsing the keywords out of the search. Which means they are intentionally loo…
It seems entirely appropriate if you have a legitimate right to use that link to parse it for structure.
Re: Matt Cutt's thoughts on Google Bing Debate
#164Earlier quoted context omitted.
You're addressing the letter of the "law" not the spirit. Do we really need a new formal standard to indicate that this is unethical? To me it's pretty plain from the standards the web has already agreed upon.
Absolutely a new standard is needed. This is the first time any major player has suggested that data from a users clickstream should be covered by robots.txt or anything like it. The comparison to a web directory is perfect - in the absence of a robots.txt Google and other crawlers think the data is absolutely fair game. There isn't yet a corresponding analogy for clickstream data so its fair game. Are you sure all o…
>You're addressing the letter of the "law" not the spirit. Do we really need a new formal standard to indicate that this is unethical? To me it's pretty plain from the ethical standards the web has already agreed upon.
Re: Matt Cutt's thoughts on Google Bing Debate
#165Earlier quoted context omitted.
Web content layout and text owned by website. Opt out with robots.txt. Clickstream owned by user. Opt out by turning off toolbar. Both should be respected.
"Clickstream owned by user. Opt out by turning off toolbar." It is not clear to me that the user completely owns the clickstream. If Google's terms of use say you cannot share their search results with anyone else, are you legally allowed to send the clickstream to Microsoft?
If Google's terms of use want to claim that a browser can't record clicks on their web pages, then perhaps browser terms of use will say they can do whatever they want with any data returned to them by the server.
Re: Matt Cutt's thoughts on Google Bing Debate
#166Let’s take that thought to its conclusion. If clicks on Google really account for only 1/1000th (or some other trivial fraction) of Microsoft’s relevancy, why not just stop using those clicks and reduce the negative coverage and perception of this? And if Microsoft is unwilling to stop incorporating Google’s clicks in Bing’s rankings, doesn’t that argue that Google’s clicks account for much more than 1/1000th of Bing…
What is being data mined is a bit more than a user broadcasting their own preference on the correct result. The user is broadcasting a URL which is selected based on two factors: - the user's preference - Google's ranking algorithm. Had Google not ranked that URL, the user wouldn't be broadcasting it. If there was some way to extract the factor of the user's preference of URLs as a signal without the factor of Google…
Re: Matt Cutt's thoughts on Google Bing Debate
#167Earlier quoted context omitted.
Absolutely a new standard is needed. This is the first time any major player has suggested that data from a users clickstream should be covered by robots.txt or anything like it. The comparison to a web directory is perfect - in the absence of a robots.txt Google and other crawlers think the data is absolutely fair game. There isn't yet a corresponding analogy for clickstream data so its fair game. Are you sure all o…
I was using "standard" in two different ways there. Let me fix that: >You're addressing the letter of the "law" not the spirit. Do we really need a new formal standard to indicate that this is unethical? To me it's pretty plain from the ethical standards the web has already agreed upon.
Re: Matt Cutt's thoughts on Google Bing Debate
#168Earlier quoted context omitted.
What do you suppose google does when they are faced with a novel query and their algorithm returns 10 equally good results as matches? At that point you might as well provide a whole bunch of users some permutation of the matches and take into account which links are clicked on the most. This is exactly what bing is doing and if you crawl the web then you will indeed end up with google's database, the only difference…
That's not what bing is doing in this case. What bing is doing is taking information from Google, as collected by users, and presenting that information to their own users.
Re: Matt Cutt's thoughts on Google Bing Debate
#169Earlier quoted context omitted.
> You can't claim ownership of facts, even if you discovered them first, or created them into existence. It's not as straightforward as all that. Suppose you write a book. The existence of that book is a fact. The words that are in that book are now a fact. Now, I create a book that says, "The following book was written by DennisM, this is a list of the words that are in that book, in order" and then precede to dupli…
Actually, it is as straightforward as all that. What's alleged is that someone searches for DenisM's book on Google, and when they click on one of the results, just that piece of data is forwarded to Microsoft. Not a scraping of all of Google's search results and nothing like the contents of the book. As to "the central idea of copyright", bear in mind that in addition to not being able to copyright ideas or data, th…
Explain this to me. In terms of raw bytes it's probably more data than every book that has ever been written. It was certainly more expensive to write than any particular book. What is the important difference here?
To make it a little more concrete, suppose someone was doing the same thing with google's streetview data. Now, it's an empirical fact that if you go to such-and-such an address, and look around, it looks like the pictures that google took. Does that mean it's ok to use those pictures, just because their resemblance to reality is factual?
Re: Matt Cutt's thoughts on Google Bing Debate
#170I'm shocked at the amount of suppport on HN (thus far) for Bing's attitude on this issue. I understand that they may well not target Google's SERPs specifically in their clickstream analysis but they should certainly have excluded Google from it, for ethical reasons. Google state that Bing created associations from clickstreams through Google's SERPs on common queries (e.g. the tarsorrhaphy spell check test), not jus…