Live data from Hacker News

The Surprising Power of Online Experiments (2017)

hbr.org

21–27 of 27 posts

Re: The Surprising Power of Online Experiments (2017)

#21
I hypothesize that many of these statistical tests have led to worse outcomes. Not in the commonly espoused bad for society externality view, but as a straight up bad for business view. There are three problems I see with A/B tests on a vast scale for small tests.

1. If the effect size is incredibly small, which most minor UI changes will be, finding statistical tests to prove them is really difficult. If you’re looking for an incredibly small positive effect, even with hundreds of thousands of sample size, the probability of rejecting the null hypothesis while the effect is actually a negative influence is surprisingly high! Very easy to make mistakes.

2) short term gains on engagement may lead to long term disengagement.

3) business incentives for management are easily misaligned. I would imagine a dominant negative influence is managers exaggerating the statistical influence found in a test because that means they get to lead the change, on an otherwise vast tech ecosystem the performance of which probably won’t change all that much. Attribution is also hard (how sure are they on how much to attribute here?) so credit is difficult to allocate beyond initial value sizing.

Re: The Surprising Power of Online Experiments (2017)

#22
post #16

Source: https://hbr.org/2017/09/the-surprising-power-of-online-exper... I used to work with Ron Kohavi in his group.

Ouch. We changed the URL from http://blog.rootshell.ir/2019/02/how-to-increase-annual-reve..., which seems to have copied that content, and banned that site.

Thanks for the heads-up.

Re: The Surprising Power of Online Experiments (2017)

#23
post #22
post #16

Source: https://hbr.org/2017/09/the-surprising-power-of-online-exper... I used to work with Ron Kohavi in his group.

Ouch. We changed the URL from http://blog.rootshell.ir/2019/02/how-to-increase-annual-reve... , which seems to have copied that content, and banned that site. Thanks for the heads-up.

Thanks. I was wondering why Ron was writing there; he wasn't.

Re: The Surprising Power of Online Experiments (2017)

#24

I hypothesize that many of these statistical tests have led to worse outcomes. Not in the commonly espoused bad for society externality view, but as a straight up bad for business view. There are three problems I see with A/B tests on a vast scale for small tests. 1. If the effect size is incredibly small, which most minor UI changes will be, finding statistical tests to prove them is really difficult. If you’re look…

I literally did a research project for a firm recently where we lowered wages. Workers are super monitored, I have data on every minute of their day. Post wage cut, workers worked just as hard. Awesome! 5 months later, the best people are (significantly) gone. 7 months later, and it is demonstrably Obvious that this was a value destroying move--average fixed effects look horrible now, yet this is only clear to those of us watching from the outside. Internally, the difference is completely overlooked.

I needed to find cites, and ironically I found the Parent article this morning. This one is way better about the lies we tell ourselves: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3204791

Re: The Surprising Power of Online Experiments (2017)

#25
post #5

Anybody have a recommendation for an A/B testing service? We've talked to Optimizely but their pricing was going to come in at the same ballpark as our AWS spend (into the six-figure range), which seems absurd. They charge based on monthly users, but a lot of our traffic consists of organic search bounces. For now we just want to run ~5 experiments per month, want to record events server-side so we can be sure not to…

You can use VWO.com

PS: it’s a product that I launched here on HN 9 years ago and wouldn’t have been possible without the awesome community.

Re: The Surprising Power of Online Experiments (2017)

#26

I hypothesize that many of these statistical tests have led to worse outcomes. Not in the commonly espoused bad for society externality view, but as a straight up bad for business view. There are three problems I see with A/B tests on a vast scale for small tests. 1. If the effect size is incredibly small, which most minor UI changes will be, finding statistical tests to prove them is really difficult. If you’re look…

1) Agreed that the smaller the effect, the more statistical power (usually from a larger sample size) you need to detect them. But to assume that all changes have tiny effects, and therefore not detectable and a waste of time, is a flawed assumption.

Once upon a time we published over 100 a/b tests here: https://www.goodui.org/evidence/ and clearly the relative effects vary (not all single changes have always a small effect).

More so, the effects of a/b tests can be further increased by grouping multiple higher confidence ideas together into a single variation.

2) Short term gains may (or may not) lead to long term disengagement. Measuring micro (shallow) and macro (deeper) metrics would be the right way to answer this.

Re: The Surprising Power of Online Experiments (2017)

#27

I hypothesize that many of these statistical tests have led to worse outcomes. Not in the commonly espoused bad for society externality view, but as a straight up bad for business view. There are three problems I see with A/B tests on a vast scale for small tests. 1. If the effect size is incredibly small, which most minor UI changes will be, finding statistical tests to prove them is really difficult. If you’re look…

I literally did a research project for a firm recently where we lowered wages. Workers are super monitored, I have data on every minute of their day. Post wage cut, workers worked just as hard. Awesome! 5 months later, the best people are (significantly) gone. 7 months later, and it is demonstrably Obvious that this was a value destroying move--average fixed effects look horrible now, yet this is only clear to those…

How do you measure people like that?
Post reply on HN