Live data from Hacker News

Challenging the Bing It On Challenge

freakonomics.com

11–20 of 63 posts

Re: Challenging the Bing It On Challenge

#11
Hmm, they used Mechanical Turk "workers" to this experiment, paying each worker 40 cents per answer, which netted them 400 responses. They upped the payout to $1 and hit 1000 soon enough. They mixed the two response sets in the final analysis.

I'm SURE that didn't skew the results at all.

They respond to this in the study, but it's going to take a lot to convince me that Mechanical Turk is a respectable source of subjects for consumer behavior research.

The real news here is that Bing and Google are virtually tied with respect to search quality. But that's been my experience based on using DuckDuckGo (which is basically a rebranded Bing afaik).

Re: Challenging the Bing It On Challenge

#12
The article is good, but the math sucks. Statistically, 1000 elements is a large enough sample for binary studies, and is consistent with the number of people that are usually asked about their voting preferences on most polls. Also, the size of the sample does not depend on the size of the population you want to sample, provided that the population is large enough.

Now, whether the sample was actually representative of the population in question is another matter.

Re: Challenging the Bing It On Challenge

#13
post #6

Its a little interesting, I tried Bing for a bit to avoid Google's onerous privacy issues, and I gave up. I moved to DuckDuckGo, haven't been back. The results are usually more relevant for me and I don't have to feed the Big-G machine.

What privacy issues does Google have that Bing doesn't?

Re: Challenging the Bing It On Challenge

#14
post #3
post #2

It's great that they did this study. I wish more people would publicly challenge many of the outlandish claims in advertising. But it seems rather hypocritical to knock Microsoft for a sample size of "nearly 1,000" when their own study "obtained 1,008 Bing It On challenge responses from the MTurk platform and narrowed our analysis to 985 respondents who submitted screen shots for 4925 searches."

Not in this case, no. They were trying to replicate Microsoft's results, which entails using a similar sample size.

I don't think that's quite right. Imagine if I flip a coin and it comes up heads, so I say "coins always come up heads". You don't flip a coin once to prove me wrong.

I think the right thing to do is to conduct a much bigger study, get a much better estimate of the preference distribution, and then derive the % probability of getting a result at least as favorable as the Bing result by pure chance.

If you can show there's only e.g. a 1% chance that the Bing result arose through random variation, then Bing definitely has some tough questions to answer. (e.g. did they run 100 different studies, and just cherry-picked the best result).

Re: Challenging the Bing It On Challenge

#15
post #3

Earlier quoted context omitted.

Not in this case, no. They were trying to replicate Microsoft's results, which entails using a similar sample size.

I don't think that's quite right. Imagine if I flip a coin and it comes up heads, so I say "coins always come up heads". You don't flip a coin once to prove me wrong. I think the right thing to do is to conduct a much bigger study, get a much better estimate of the preference distribution, and then derive the % probability of getting a result at least as favorable as the Bing result by pure chance. If you can show th…

Yes, that would have been a great way to do it as well. However, the most interesting aspect of the study to me was that they used Microsoft's methodology, comparable sample sizes, and examined the effect of changing search terms from Microsoft's "recommended" list to ones that study participants suggested themselves.

Incidentally, you don't need a giant sample to estimate the variance on the Google/Bing ratio -- you can use any number of resampling techniques like bootstrapping to get that estimate.

Re: Challenging the Bing It On Challenge

#18
post #8

Marketing aside I find Bing results good. Google is kind of turning into singularity and tends to return only google (or US) centered results (for example youtube only videos). Bing seems to have better diversity.

The opposite is just as annoying if not more:

I search in English and google insists on translating my search terms to the local language and displaying local search results.

Again: idea for a startup: -like google but 5 years ago.

Re: Challenging the Bing It On Challenge

#19
post #15

Earlier quoted context omitted.

I don't think that's quite right. Imagine if I flip a coin and it comes up heads, so I say "coins always come up heads". You don't flip a coin once to prove me wrong. I think the right thing to do is to conduct a much bigger study, get a much better estimate of the preference distribution, and then derive the % probability of getting a result at least as favorable as the Bing result by pure chance. If you can show th…

Yes, that would have been a great way to do it as well. However, the most interesting aspect of the study to me was that they used Microsoft's methodology, comparable sample sizes, and examined the effect of changing search terms from Microsoft's "recommended" list to ones that study participants suggested themselves. Incidentally, you don't need a giant sample to estimate the variance on the Google/Bing ratio -- you…

Good point - particularly as it's a fairly simple universe of answers. I wonder if we could get the raw data. But, I do think the Mechanical Turk vs "shopping mall" selection-bias factor overwhelms everything else.

Re: Challenging the Bing It On Challenge

#20
Bing's abysmal performance on WP7 has seriously hurt the brand for me. If you're going to bundle the search engine into a dedicated button on your phone, you should probably avoid making it awful.
Post reply on HN