Live data from Hacker News

Conservation of Intent: why A/B tests aren’t as effective as they look

andrewchen.co

1–10 of 59 posts

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#3
I could not disagree with this more. I remember vividly having this "low-intent" vs "high-intent" debate at Thumbtack, when we rolled out changes that A/B tests showed increased conversion (by a lot), but some people in the company thought the changes were ugly and "off-brand" and argued they brought in the wrong type of customers. So we ran the test again that we knew raised conversion by a lot, and then followed the 2 cohorts of customers and watched their behavior. The control group vs the 10% more from whatever the test was that increased conversion. They behaved exactly the same. They came back again at the same rates. They made the same amount of profit (per customer). Their response rates to emails were the same. They closed jobs at the same rates. As far as we could tell they were identical.

I have to admit I was a little surprised too, but for our business it didn't seem this "high-intent" vs "low-intent" distinction existed. And with that out of the way we continued to optimize conversion rates, and our revenue continued to go up.

Every company is different so I don't want to generalize too much, but if somebody tells me they ran an A/B test that said some key flow went up 10%, but then afterwards the traffic/revenue/whatever didn't go up 10%, I think the most likely candidate is bad test design. Humans are really good at rigging A/B tests to produce wrong results in their favor. I guarantee every single company who isn't maniacal about A/B testing does at least one of the following:

- Uses a tool to grade A/B tests that isn't statistically sound

- Let's people check tests too often and allows them to stop the test when it hits a good result

- Running a test with a lot of similar variations and cherry picking the best one

- Doesn't plan for enough traffic to detect the percentage of change their test is likely to produce

All of these create the potential for the perceived gains of the A/B test not matching up with real world result.

I'm not saying the distinction between "low-intent" and "high-intent" customers doesn't exist, but it is fairly easy to test for. Do that test for your business and see if that distinction exists. But don't use it as some magical explanation for why your A/B tests aren't producing the results you want as this article suggests.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#4
Wait, so you’re telling me the laziest form of scientific analysis, the A/B test, doesn’t produce accurate results? Colour me shocked.

A/B tests routinely leave out important observations, have way too small a scope, uncontrolled populations, I could go on... they run the gamut of anti-patterns.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#5
post #3

I could not disagree with this more. I remember vividly having this "low-intent" vs "high-intent" debate at Thumbtack, when we rolled out changes that A/B tests showed increased conversion (by a lot), but some people in the company thought the changes were ugly and "off-brand" and argued they brought in the wrong type of customers. So we ran the test again that we knew raised conversion by a lot, and then followed th…

> Let's people check test too often and allows them to stop the test when it hits a good result

I admit to attempting to be guilty of this in the past and being stopped by our analytics team (in the sense that they took the time to patiently explain to me why what I was doing was statistically unsound). It's not obvious, IMO.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#6
The title is misleading relative to the article's content. Surely, as the author points out, sometimes A/B tests can be misleading especially if you ignore longer term cohort analysis, etc.

But often times, if you fix an obviously broken part of your funnel, particularly in the early acquisition stages, you're fixing things that are universally lifting the amount of people who ultimately are able to engage with your brand and product to the point where they can even form intent. The reality is most people are only willing to give you a tiny bit of their time during their first one or two engagements with your brand, so at that stage you're trying to sell them on your product, and build intent. A/B testing helps reduce the friction needed to get them through the core of your sales pitch.

It's easy to come up with a thought experiment that shows A/B testing can sometimes be as simple as you'd imagine: just break the site. Your conversion drops to 0%, now split test the fix. Like magic, your control stays at 0% and your variant returns to normal. Nothing about "intent" in this scenario, this is pure friction resolution. Just a thought experiment, but shows that surely there are plenty of places where pure A/B testing and removing friction is a net positive without any fretting over this "conservation of intent" issue.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#7
post #4

Wait, so you’re telling me the laziest form of scientific analysis, the A/B test, doesn’t produce accurate results? Colour me shocked. A/B tests routinely leave out important observations, have way too small a scope, uncontrolled populations, I could go on... they run the gamut of anti-patterns.

....and criticising statistics is the laziest kind of scientific criticism....

A/B tests are fine. They work. They allow inference of causality. They are easy to understand, and can be fun to run. They get you 90% of wherever you want to go, and such over-the-top criticism just seems like badly executed pretentiousness.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#8
post #4

Wait, so you’re telling me the laziest form of scientific analysis, the A/B test, doesn’t produce accurate results? Colour me shocked. A/B tests routinely leave out important observations, have way too small a scope, uncontrolled populations, I could go on... they run the gamut of anti-patterns.

This. It has always seemed pretty nuts to me the slippage between the sort of idealized multi-armed bandit mechanics of A/B testing (the theoretically sound basis) and the actual real-world situations with enormous hypothesis spaces and gnarly sampling problems. But I guess even finding local minima / maxima is useful?

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#9
post #7
post #4

Wait, so you’re telling me the laziest form of scientific analysis, the A/B test, doesn’t produce accurate results? Colour me shocked. A/B tests routinely leave out important observations, have way too small a scope, uncontrolled populations, I could go on... they run the gamut of anti-patterns.

....and criticising statistics is the laziest kind of scientific criticism.... A/B tests are fine. They work. They allow inference of causality. They are easy to understand, and can be fun to run. They get you 90% of wherever you want to go, and such over-the-top criticism just seems like badly executed pretentiousness.

You’re partly right - a scientific endeavor to figure out the color of a button would be over-the-top, because it’s not that important.

But to the article’s point, if you’re running banking software or something, your users don’t give a shit what the button colours are; they will slog through whatever you develop because they need to get stuff done.

A/B tests are a small tool that sometimes get taken too far or used in the wrong context.

Post reply on HN