Every time they say "classic statistics" just insert "what we did before now" and see how frustrated you get with this announcement. The whole point of using them is that people don't need a statistician because the tool should make it easy to run solid tests. That of course hasn't been the case and they're finally admitting it.
Thanks for your comment. This is Darwish, the Product Manager working on Stats Engine. You are correct "classic statistics" is the method we used in the past. It also what is most commonly used in industry (the main reason we started with this method). This was not an easy project for us to take on, but after talking to customers and looking at our historical experiment data, it was clear how important this problem w…
Optimizely Statistics Engine
41–50 of 57 posts
Re: Optimizely Statistics Engine
#42Based on my admittedly limited understanding of stats, unless you set the sample size and decide what significance is in advance your test will probably misinform you. Nothing on this landing page explains to me how this new thing might mean otherwise and it really doesn't help that the page is otherwise full of hubris, eg:"goodbye traditional statistics". Somehow it seems unlikely that a web startup just invalidated all of statistics
Re: Optimizely Statistics Engine
#43I wrote about the problem with sequential testing in online experiment three years ago on the Custora blog [1]. And Evan Miller wrote about it two years before me on his blog [2]. I'm glad to see Optimizely finally getting on board. Communicating statistical significance to marketers is always challenging, and I'm sure this will lead to better decisions being made. [1] http://blog.custora.com/2012/05/a-bayesian-appro…
Answering addressing a few comments right here. I think the industry deserves a lot of credit in its efforts to help those wanting to run A/B tests. Many people were aware these were issues and many actually tried to fix it (us included). There are many blog posts in the community about why continuous monitoring is dangerous, why you should use a sample size calculator, how to properly set a Minimum Detectable Effect etc... We were part (and definitely not the first) of this group as we published a sample size calculator and spent a lot of time working with our clients on running tests with a safe testing procedure.
However, after doing this and looking more closely attempting to quantify the effect of these efforts we saw an opportunity for a simpler solution that could help even more people. Sequential Testing was this solution, and it's had success in other applications. We wanted to bring sequential testing to A/B testing and take the hard work out of doing it correctly. Specifically, we have built on that groundwork laid in 50's and 60's by providing an always valid notion of p-value that customers are looking for.
While traditional sequential testing combats the continuous monitoring problem well, they require you to have an intimate understanding of the solution that can pose cognitive hurdles for those not well-versed in statistics. You have to either know your target effect size, or have in mind a maximum allowable number of visitors and understand how changes in these will affect the run time of your test. What’s more, it is not straightforward to translate results to standard measures of significance such as p-values. This is actually where the biggest research contribution of Stats Engine comes in. We allow you to run a test, detect a range of effect sizes and provide an always valid FDR-adjusted p-value as opposed to a set of stopping rules that bounds Type 1 error at say 5%. The error rates are valid no matter how the user chooses to interact with the A/B test. Also, FDR control itself has only been around over the last 20-25 years.
Our biggest industry contribution is probably much simpler in us moving a lot of the market to sequential testing more generally. We are happy to be in the position to help build on research and bring this to practical applications.
Re: Optimizely Statistics Engine
#44Earlier quoted context omitted.
Thanks for your comment. This is Darwish, the Product Manager working on Stats Engine. You are correct "classic statistics" is the method we used in the past. It also what is most commonly used in industry (the main reason we started with this method). This was not an easy project for us to take on, but after talking to customers and looking at our historical experiment data, it was clear how important this problem w…
What's so particularly embarrassing is that you clearly did not have any competent statisticians on board until now. This was not some big surprise that needed "a lot of resources to fix." This is something that should be obvious to anyone who understands hypothesis testing, and is something that statisticians have been describing how to do correctly for over 50 years: http://www.ncbi.nlm.nih.gov/pmc/articles/PMC1551…
Re: Optimizely Statistics Engine
#45Re: Optimizely Statistics Engine
#46Earlier quoted context omitted.
Hello, Leo, Optimizely's in-house statistician here. The graph you reference is a schematic to show the differences between Optimizely’s previous statistical platform and Stats Engine. It shows a monotone non-decreasing significance because under our sequential testing framework, the significance value represents the total amount of accumulated evidence against the null hypothesis of no difference between a variation…
> It shows a monotone non-decreasing significance because [the value] represents the total amount of accumulated evidence against the null hypothesis. > if we instead looked continuously at a classical t-test, is the significance would oscillate near the significance threshold So there's your answer: the y-axis on the chart has an unlabeled different meaning for the blue line. While I have you here Leo, can you expla…
The amount of accumulated evidence for X is exactly a p-value, or a measurement which can tell you if there is enough evidence in the experiment to contradict a hypothesis of “no difference between a baseline and variation.” A high p-value, or low significance tells you there is a lack of evidence to make this claim.
You bring up a very interesting point which is that with sequential testing it is actually possible to also look for evidence of ‘not X’ or that there really is no detectable difference. This works by ‘flipping the hypothesis test on it’s head’ and allows for a mathematical formulation of stopping early for futility. We do not currently offer this in Stats Engine because we believe it’s the less important quantity of the two, but it may be the focus of future research.
Re: Optimizely Statistics Engine
#47Earlier quoted context omitted.
Hello, Leo, Optimizely's in-house statistician here. The graph you reference is a schematic to show the differences between Optimizely’s previous statistical platform and Stats Engine. It shows a monotone non-decreasing significance because under our sequential testing framework, the significance value represents the total amount of accumulated evidence against the null hypothesis of no difference between a variation…
> It shows a monotone non-decreasing significance because [the value] represents the total amount of accumulated evidence against the null hypothesis. > if we instead looked continuously at a classical t-test, is the significance would oscillate near the significance threshold So there's your answer: the y-axis on the chart has an unlabeled different meaning for the blue line. While I have you here Leo, can you expla…
Re: Optimizely Statistics Engine
#48Earlier quoted context omitted.
Thanks for your comment. This is Darwish, the Product Manager working on Stats Engine. You are correct "classic statistics" is the method we used in the past. It also what is most commonly used in industry (the main reason we started with this method). This was not an easy project for us to take on, but after talking to customers and looking at our historical experiment data, it was clear how important this problem w…
This feels like a very wasteful approach to me. I understand the need to protect against the variation of p-value with each test. However, now data needs to mature in two places. Bayesian testing is a natural cyclical approach where past information coupled with data generates new beliefs which become past information. Most conversion rates are small, say less than 30%. Hence, their differences are smaller (as we mov…
Second, I want to point out that Stats Engine is not a Bayesian test. We do not recompute a posterior from past information after every visitor and use this directly to get significance. Instead such calculations are used as inputs to determine how much information we have compared to a situation of zero effect size. There’s only ‘one lense’ still because we use all this to make and guarantee the usual Frequentist hypothesis testing statements, but now factoring in that an experimenter can look at results at any time.
Re: Optimizely Statistics Engine
#49I am surprised by all the negative commentary here. On the whole, companies like Optimizely, RJMetrics, Custora, and others are doing more to push statistical analysis to the mass market than anyone else. These tools are not designed for statisticians or ML practitioners so it makes sense they do not put language like Bayesian, etc. front and center. IMO, the more people using data to make decisions, the better.
It's not that they don't put in 'language like Bayesian', it's a different method. Yes, it is an improvement on the t-test straw-man they mention, but it's less flexible and powerful than Bayesian methods. Once you have a posterior, you can ask different questions that their p-values/confidence intervals don't address. For example, probability of an x% increase in conversion rate, or the risk associated with choosing…
I agree that a benefit of Bayesian analysis is flexibility. Different posterior results are possible with different priors. But in practice this can be a hindrance as well as a benefit. When answer depend on a choice of prior, misusing, or misunderstanding the prior can lead to incorrect conclusions.
There is also a very attractive feature of Frequentist guarantees specifically for A/B testing. They make statements on the long-run average lift, which is a quantity that many businesses care about: what will my average lift be if I implement a variation after my A/B test?
That said, we have, and continue to look at Bayesian methods because we don’t feel that we have to be in either a Frequentist or Bayesian framework, but rather use the tools that are best suited to answer the sorts of statistical questions our customers encounter.
Finally, there have been some very interesting results lately on the connections between sequential testing and bandits! (for example, see here: http://auduno.github.io/SeGLiR/documentation/reference.html )
Re: Optimizely Statistics Engine
#50Earlier quoted context omitted.
This feels like a very wasteful approach to me. I understand the need to protect against the variation of p-value with each test. However, now data needs to mature in two places. Bayesian testing is a natural cyclical approach where past information coupled with data generates new beliefs which become past information. Most conversion rates are small, say less than 30%. Hence, their differences are smaller (as we mov…
I do agree with you that with sequential testing it is possible to get much slower results. This is actually similar to using a sample size calculator for a classical t-test. If you set your minimum detectible effect (MDE) much smaller than the actual effect size of your A/B test, you will end up waiting for many more visitors than were needed to detect significance. We have looked at many historical A/B tests at Opt…