Live data from Hacker News

Why Wikipedia's A/B testing is wrong

synference.blogspot.com

1–10 of 31 posts

Re: Why Wikipedia's A/B testing is wrong

#3
post #2

"Highly successful" and "All wrong"? There's an oxymoron if ever I've heard one.

(Synference cofounder here)

It has successfully increased Wikipedia's donation revenue; however, some fundamental assumptions baked into the AB-testing approach are all wrong; this leaves a lot of money on the table.

Instead of trying to find a single best version for everyone, people should realise that different segments of their users are going to have different preferences. Machine learning can find these in an automated way, and there's a lot of value to be gained - but people need to see past the standard 'one size fits all' AB-testing approach to take advantage of it.

Re: Why Wikipedia's A/B testing is wrong

#4
Really interesting concept, but what other information can you grab from a user who has never visited your site before?

There's the example of location, and I guess you can also get browser information (mobile device vs. desktop), but other than that, I feel there isn't that much in terms of context to work with with first-time anonymous users, which is the most common case for A/B testing.

Would love to hear how else you're classifying users, and what other buckets you're using.

Re: Why Wikipedia's A/B testing is wrong

#5
It's awesome to see a Bandit-based algorithm being used. I think this is really similar to what is called Adaptive Sampling (also based on the Bandit problem), and it's been used in multiple contexts. My favorite has been in clinical trials where adaptive sampling was able to save lives.

Read more about the research done in Clinical Trials, and one of my favorite Professors here: http://web.eecs.umich.edu/~qstout/AdaptiveDesign.html

Re: Why Wikipedia's A/B testing is wrong

#6
post #3
post #2

"Highly successful" and "All wrong"? There's an oxymoron if ever I've heard one.

(Synference cofounder here) It has successfully increased Wikipedia's donation revenue; however, some fundamental assumptions baked into the AB-testing approach are all wrong; this leaves a lot of money on the table. Instead of trying to find a single best version for everyone, people should realise that different segments of their users are going to have different preferences. Machine learning can find these in an a…

There's a huge difference between "all wrong" and "has some fundamental assumptions that could lead to non-optimal results if you wanted to push the results even further". The former implies, upon initial reading of the headline, that whatever Wikipedia is doing doesn't even give you a better solution.

A better headline might have been "Hidden assumptions in Wikipedia's A/B testing and how it could be improved". Also, why are we picking on Wikipedia here? Don't a lot of companies claim to do A/B testing and isn't this a problem inherent in all simple A/B testing?

Re: Why Wikipedia's A/B testing is wrong

#9
post #6
post #3

Earlier quoted context omitted.

(Synference cofounder here) It has successfully increased Wikipedia's donation revenue; however, some fundamental assumptions baked into the AB-testing approach are all wrong; this leaves a lot of money on the table. Instead of trying to find a single best version for everyone, people should realise that different segments of their users are going to have different preferences. Machine learning can find these in an a…

There's a huge difference between "all wrong" and "has some fundamental assumptions that could lead to non-optimal results if you wanted to push the results even further". The former implies, upon initial reading of the headline, that whatever Wikipedia is doing doesn't even give you a better solution. A better headline might have been "Hidden assumptions in Wikipedia's A/B testing and how it could be improved". Also…

AB-testing is much better than doing nothing.

But you have to be careful with any powerful tool, in case its success blinds you to its weaknesses. "When all you have is a hammer, everything starts to look like a nail". We think that's happening with AB-testing at the moment.

We think Wikipedia is awesome, and would love them to get more donations by using a more sophisticated approach.

Yes, most companies doing AB-testing, if they have the ability to personalise their user experience (i.e. they aren't trying to quickly find the single best UI) could benefit.

However, Wikipedia is a really good example to start with - the phrase 'Wikipedia needs those nickels' is a great example - resonates well with US donators, will probably work in Canada, but what about the UK? Australia?

It's obvious once its pointed out - but wouldn't it be better if the system automatically realises this? And considers all the combinations? That's our point.

Re: Why Wikipedia's A/B testing is wrong

#10
This is a little silly. Of course every form would convert better if you perfectly customized it to the current user, no one really disagrees with that. But it doesn't mean it's always practical or worth the effort, or that the Wikipedia team have the time to tackle that at this moment.
Post reply on HN