Live data from Hacker News

Conservation of Intent: why A/B tests aren’t as effective as they look

andrewchen.co

21–30 of 59 posts

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#21
post #5
post #3

I could not disagree with this more. I remember vividly having this "low-intent" vs "high-intent" debate at Thumbtack, when we rolled out changes that A/B tests showed increased conversion (by a lot), but some people in the company thought the changes were ugly and "off-brand" and argued they brought in the wrong type of customers. So we ran the test again that we knew raised conversion by a lot, and then followed th…

> Let's people check test too often and allows them to stop the test when it hits a good result I admit to attempting to be guilty of this in the past and being stopped by our analytics team (in the sense that they took the time to patiently explain to me why what I was doing was statistically unsound). It's not obvious, IMO.

Well, you can arrange the test so checking often is not a problem. It may even be the optimal way.

But yes, statistics is not obvious at all.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#22
post #3

I could not disagree with this more. I remember vividly having this "low-intent" vs "high-intent" debate at Thumbtack, when we rolled out changes that A/B tests showed increased conversion (by a lot), but some people in the company thought the changes were ugly and "off-brand" and argued they brought in the wrong type of customers. So we ran the test again that we knew raised conversion by a lot, and then followed th…

Did you read the article? The real point has less to do with intent that you realize.

Most of the discussion has to do with having healthy skepticism for a vendor making exaggerated claims; and ensuring that you retain customers.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#23
post #4

Wait, so you’re telling me the laziest form of scientific analysis, the A/B test, doesn’t produce accurate results? Colour me shocked. A/B tests routinely leave out important observations, have way too small a scope, uncontrolled populations, I could go on... they run the gamut of anti-patterns.

What the software industry calls an "A/B test" is what scientists call a "randomized controlled trial",* and it's generally considered to be the best type of experiment you can do.

The fact that people routinely screw up their experiments in the ways you list and others isn't an indictment of the A/B test methodology.

* Well actually the term gets used to mean a lot of things but insofar as you can give it a meaningful definition, it means randomized controlled trial.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#24
post #3

I could not disagree with this more. I remember vividly having this "low-intent" vs "high-intent" debate at Thumbtack, when we rolled out changes that A/B tests showed increased conversion (by a lot), but some people in the company thought the changes were ugly and "off-brand" and argued they brought in the wrong type of customers. So we ran the test again that we knew raised conversion by a lot, and then followed th…

So damn true, it's so difficult to get statistically significant A/B test results, and the popular tools actively lead you astray. It's rare enough for a startup to have the traffic to meaningfully get results on their homepage before the heat death of the universe, let alone random landing or in-app pages. I'd recommend anyone reading this who does or wants to do A/B tests read: https://www.evanmiller.org/how-not-to-run-an-ab-test.html It was one of the trickiest lessons I had to learn. Once you realize you need to set your sample size in advance, you actually have to do the math to figure out the traffic you'll need. That's when reality hits you.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#25
post #9
post #7

Earlier quoted context omitted.

....and criticising statistics is the laziest kind of scientific criticism.... A/B tests are fine. They work. They allow inference of causality. They are easy to understand, and can be fun to run. They get you 90% of wherever you want to go, and such over-the-top criticism just seems like badly executed pretentiousness.

You’re partly right - a scientific endeavor to figure out the color of a button would be over-the-top, because it’s not that important. But to the article’s point, if you’re running banking software or something, your users don’t give a shit what the button colours are; they will slog through whatever you develop because they need to get stuff done. A/B tests are a small tool that sometimes get taken too far or used…

> a scientific endeavor to figure out the color of a button would be over-the-top, because it’s not that important.

Not saying A/B testing shades of blue like Google reportedly does is anything useful, but picking a different button color altogether reportedly can make a difference.

Even if it's 1% more sign-ups each month, that can translate to a few more potential clients per month - and more word of mouth. If you think that is too trivial a difference to matter, think about what 1% more interest would mean over your career as you compound interest for your retirement.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#26

A/B tests tell you about short term gains, but don't tell you about long term issues you may be accumulating due to things like dark patterns, clickbait headlines, shoddy article topics and more. A/B tests don't take into account the loss of prestige or reputation that the options give. I've seen this repeatedly with ArsTechnica, which has devolved into so much political and clickbait material that I don't even reall…

> I've seen this repeatedly with ArsTechnica

Can you expand? I am an Ars reader, and I too find it frustrating compared to what it used to be. I wish there were more in-depth technical articles; I find it too light on details, written for a non-tech audience. I really, really wish it had solid technical content, since it's what got me reading it.

But I don't characterise it as click-bait (it seems clear) nor as especially political, except in so far as politics intersects technology and climate. In both those it seems balanced or erring towards freedom.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#27

A/B tests tell you about short term gains, but don't tell you about long term issues you may be accumulating due to things like dark patterns, clickbait headlines, shoddy article topics and more. A/B tests don't take into account the loss of prestige or reputation that the options give. I've seen this repeatedly with ArsTechnica, which has devolved into so much political and clickbait material that I don't even reall…

So true about the danger of focusing on short term gains. I think that it is a problem of picking the wrong things to optimize and then sampling on the wrong entities. Conversion rate is often the most commonly used target, but, it is short sighted. In the big picture a company is really just trying to maximize profit. Note that you can't use traditional A/B testing tools on a value like total profit. You need an average if you are going to run an A/B test. For conversion rates your average is conversions per user. In this case you sample on the denominator -- the users. What do you sample on if you are testing total profit per quarter? This is the ultimate conundrum for websites who use A/B testing. They don't understand how they might actually get more new/returning users by simply improving the user experience. Instead, they run marketing campaigns to drive growth and then squeeze every dollar out of the customers when they arrive. This is not a good long term strategy. I have reviewed hundreds of A/B tests over the years. The only ones that I think succeeded were the ones that focused on improving the users experience.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#28
post #24
post #3

I could not disagree with this more. I remember vividly having this "low-intent" vs "high-intent" debate at Thumbtack, when we rolled out changes that A/B tests showed increased conversion (by a lot), but some people in the company thought the changes were ugly and "off-brand" and argued they brought in the wrong type of customers. So we ran the test again that we knew raised conversion by a lot, and then followed th…

So damn true, it's so difficult to get statistically significant A/B test results, and the popular tools actively lead you astray. It's rare enough for a startup to have the traffic to meaningfully get results on their homepage before the heat death of the universe, let alone random landing or in-app pages. I'd recommend anyone reading this who does or wants to do A/B tests read: https://www.evanmiller.org/how-not-to…

Yup. The best thing for running an A/B test you can do, by far, is setting up the rules in advance. Our protocol was something like this:

This test is going to have 2 variations, looking for a 5% increase of the conversion rate that is currently X%, and to do so the test will run for Y iterations (based on the company standards for significance and power).

If the test shows a >= 5% increase, it wins and we use it. Yay! If not, we assume it is no different than the baseline, record the results and discard it. You are welcome to peek at the results all you want, but no tests are stopped early and no decisions are made until it reaches the set amount of iterations. This isn't the only good or valid A/B testing protocol, but it does force people to consider the costs of their tests in advance (in terms of time and iterations required), which I think had a positive effect on the type of tests people ran.

Just having the calculator discourages people from running tests with tons of variations looking for tiny increases (like the famous try different colors of your submit button), because just looking at the iterations required by the calculator it becomes obvious to everybody that those types of tests just cannot show any meaningful results for your average website.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#29
post #5
post #3

I could not disagree with this more. I remember vividly having this "low-intent" vs "high-intent" debate at Thumbtack, when we rolled out changes that A/B tests showed increased conversion (by a lot), but some people in the company thought the changes were ugly and "off-brand" and argued they brought in the wrong type of customers. So we ran the test again that we knew raised conversion by a lot, and then followed th…

> Let's people check test too often and allows them to stop the test when it hits a good result I admit to attempting to be guilty of this in the past and being stopped by our analytics team (in the sense that they took the time to patiently explain to me why what I was doing was statistically unsound). It's not obvious, IMO.

Meh.

There is statistics for the purpose of uncovering Truth, and statistics for the purpose of making a business decision. The difference is that when we talk about Truth, a small error is still an error. When we make business decisions, it is fine to make a decision that is probably right, and we know isn't far wrong.

Here is a perfectly valid test procedure that illustrates the difference. Decide the most time you would be willing to spend to get a test result. Multiply that by current conversion rates to get N, the number of conversions that you expect to see by the end of the test.

Start running the test with two variations. Stop at any point if one variation is at least sqrt(N) conversions ahead of the other. Stop at N if there is no clear winner and go with whoever is ahead, even by a hair.

Here are features of this test procedure.

o You always make a decision.

o Running a test has a known fixed cost. You know how long it takes. And a bad idea will cost you no more than sqrt(N) conversions to test.

o The results are very simple and easy to understand.

o Your answers are usually right.

o Your bad decisions are not very bad. If the true conversion rate for one version is better by 1/sqrt(N), you've got a 95% chance of making the right choice. You will probably never make a mistake as big as 2/sqrt(N).

The result is a test procedure that is a horrible approach for doing science, but an excellent tool for improving a business. You'll never find it in a statistics class. And I'm sure it would horrify your analytics team.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#30

A/B tests tell you about short term gains, but don't tell you about long term issues you may be accumulating due to things like dark patterns, clickbait headlines, shoddy article topics and more. A/B tests don't take into account the loss of prestige or reputation that the options give. I've seen this repeatedly with ArsTechnica, which has devolved into so much political and clickbait material that I don't even reall…

Winning the game but losing the meta-game.
Post reply on HN