Why Wikipedia's A/B testing is wrong
synference.blogspot.com
Why Wikipedia's A/B testing is wrong
1–10 of 31 posts
Re: Why Wikipedia's A/B testing is wrong
#2Re: Why Wikipedia's A/B testing is wrong
#3"Highly successful" and "All wrong"? There's an oxymoron if ever I've heard one.
It has successfully increased Wikipedia's donation revenue; however, some fundamental assumptions baked into the AB-testing approach are all wrong; this leaves a lot of money on the table.
Instead of trying to find a single best version for everyone, people should realise that different segments of their users are going to have different preferences. Machine learning can find these in an automated way, and there's a lot of value to be gained - but people need to see past the standard 'one size fits all' AB-testing approach to take advantage of it.
Re: Why Wikipedia's A/B testing is wrong
#4There's the example of location, and I guess you can also get browser information (mobile device vs. desktop), but other than that, I feel there isn't that much in terms of context to work with with first-time anonymous users, which is the most common case for A/B testing.
Would love to hear how else you're classifying users, and what other buckets you're using.
Re: Why Wikipedia's A/B testing is wrong
#5Read more about the research done in Clinical Trials, and one of my favorite Professors here: http://web.eecs.umich.edu/~qstout/AdaptiveDesign.html
Re: Why Wikipedia's A/B testing is wrong
#6"Highly successful" and "All wrong"? There's an oxymoron if ever I've heard one.
(Synference cofounder here) It has successfully increased Wikipedia's donation revenue; however, some fundamental assumptions baked into the AB-testing approach are all wrong; this leaves a lot of money on the table. Instead of trying to find a single best version for everyone, people should realise that different segments of their users are going to have different preferences. Machine learning can find these in an a…
A better headline might have been "Hidden assumptions in Wikipedia's A/B testing and how it could be improved". Also, why are we picking on Wikipedia here? Don't a lot of companies claim to do A/B testing and isn't this a problem inherent in all simple A/B testing?
Re: Why Wikipedia's A/B testing is wrong
#7Re: Why Wikipedia's A/B testing is wrong
#8"Highly successful" and "All wrong"? There's an oxymoron if ever I've heard one.
Re: Why Wikipedia's A/B testing is wrong
#9Earlier quoted context omitted.
(Synference cofounder here) It has successfully increased Wikipedia's donation revenue; however, some fundamental assumptions baked into the AB-testing approach are all wrong; this leaves a lot of money on the table. Instead of trying to find a single best version for everyone, people should realise that different segments of their users are going to have different preferences. Machine learning can find these in an a…
There's a huge difference between "all wrong" and "has some fundamental assumptions that could lead to non-optimal results if you wanted to push the results even further". The former implies, upon initial reading of the headline, that whatever Wikipedia is doing doesn't even give you a better solution. A better headline might have been "Hidden assumptions in Wikipedia's A/B testing and how it could be improved". Also…
But you have to be careful with any powerful tool, in case its success blinds you to its weaknesses. "When all you have is a hammer, everything starts to look like a nail". We think that's happening with AB-testing at the moment.
We think Wikipedia is awesome, and would love them to get more donations by using a more sophisticated approach.
Yes, most companies doing AB-testing, if they have the ability to personalise their user experience (i.e. they aren't trying to quickly find the single best UI) could benefit.
However, Wikipedia is a really good example to start with - the phrase 'Wikipedia needs those nickels' is a great example - resonates well with US donators, will probably work in Canada, but what about the UK? Australia?
It's obvious once its pointed out - but wouldn't it be better if the system automatically realises this? And considers all the combinations? That's our point.