I don't understand this paragraph. They only look for indications that the drug is better... than what?
Winning A/B results were not translating into improved user acquisition
41–50 of 67 posts
Re: Winning A/B results were not translating into improved user acquisition
#42Of course, if you're a startup, building an A/B testing tool is your last priority, so you would use an existing solution.
Are there much more advanced 'out-of-the-box' tools for testing out there besides the usual suspects, i.e. Optimizely, Monetate, VWO, etc.?
Re: Winning A/B results were not translating into improved user acquisition
#43The red flag here for me was that Optimizely encourages you to stop the test as soon as it "reaches significance." You shouldn't do that. What you should do is precalculate a sample size based on the statistical power you need, which involves determining your tolerance for the probability of making an error and on the minimum effect size you need to detect. Then, you run the test to completion and crunch the numbers…
Re: Winning A/B results were not translating into improved user acquisition
#44Oh great, another misuse of A/B testing Here's the thing, stop A/Bing every little thing (and/or "just because") and you'll get more significant results. Do you think the true success of something is due to A/B testing? A/B testing is optimizing, not archtecting.
Re: Winning A/B results were not translating into improved user acquisition
#45Peculiar use of the word bug in this context: "They make it easy to catch the A/B testing bug..."
Re: Winning A/B results were not translating into improved user acquisition
#46It seems a mod (?) changed it to "Winning A/B results were not translating into improved user acquisition".
I've seen a descriptive title left by the submitter change back to the less descriptive original by a mod. But I'm curious why a mod would editorialize certain titles and change them away from their original, but undo the editorializing of others and change them to the less descriptive originals.
Re: Winning A/B results were not translating into improved user acquisition
#47The red flag here for me was that Optimizely encourages you to stop the test as soon as it "reaches significance." You shouldn't do that. What you should do is precalculate a sample size based on the statistical power you need, which involves determining your tolerance for the probability of making an error and on the minimum effect size you need to detect. Then, you run the test to completion and crunch the numbers…
#1 - “Optimizely encourages you to stop the test as soon as it reaches ‘statistical significance.’” - This actually isn’t true. We recommend you calculate your sample size before you start your test using a statistical significance calculator and waiting until you reach that sample size before stopping your test. We wrote a detailed article about how long to run a test, here: https://help.optimizely.com/hc/en-us/articles/200133789-How-...
We also have a sample size calculator you can use, here: https://www.optimizely.com/resources/sample-size-calculator
#2 - Optimizely uses a one-tailed test, rather than a 2-tailed test. - This is a point the article makes and it came up in our customer community a few weeks ago. One of our statisticians wrote a detailed reply, and here’s the TL;DR:
- Optimizely actually uses two 1-tailed tests, not one.
- There is no mathematical difference between a 2-tailed test at 95% confidence and two 1-tailed tests at 97.5% confidence.
- There is a difference in the way you describe error, and we believe we define error in a way that is most natural within the context of A/B testing.
- You can achieve the same result as a 2-tailed test at 95% confidence in Optimizely by requiring the Chance to Beat Baseline to exceed 97.5%.
- We’re working on some exciting enhancements to our methodologies to make results even easier to interpret and more meaningfully actionable for those with no formal Statistics background. Stay tuned!
Here’s the full response if you’re interested in reading more: http://community.optimizely.com/t5/Strategy-Culture/Let-s-ta...
Overall I think it’s great that we’re having this conversation in a public forum because it draws attention to the fact that statistics matter in interpreting test results accurately. All too often, I see people running A/B tests without thinking about how to ensure their results are statistically valid.
Dan
Re: Winning A/B results were not translating into improved user acquisition
#48Earlier quoted context omitted.
More precisely, before you start the test you need to choose a "default" choice. If the default choice is the old version, then it's safe to switch to the new version provided it isn't worse. Apply the converse if your default choice is the new version. The key point here is that you aren't choosing a testing procedure , you are choosing a decision procedure .
Frequentism rears its ugly head again...
Re: Winning A/B results were not translating into improved user acquisition
#49This title used to read "How Optimizely (Almost) Got Me Fired", which is the actual title of the article. It seems a mod (?) changed it to "Winning A/B results were not translating into improved user acquisition". I've seen a descriptive title left by the submitter change back to the less descriptive original by a mod. But I'm curious why a mod would editorialize certain titles and change them away from their origina…
Re: Winning A/B results were not translating into improved user acquisition
#50The red flag here for me was that Optimizely encourages you to stop the test as soon as it "reaches significance." You shouldn't do that. What you should do is precalculate a sample size based on the statistical power you need, which involves determining your tolerance for the probability of making an error and on the minimum effect size you need to detect. Then, you run the test to completion and crunch the numbers…
For a second I thought you were Evan Miller who wrote about the exact same thing: http://www.evanmiller.org/how-not-to-run-an-ab-test.html