Live data from Hacker News

Conservation of Intent: why A/B tests aren’t as effective as they look

andrewchen.co

51–59 of 59 posts

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#51
post #5
post #3

I could not disagree with this more. I remember vividly having this "low-intent" vs "high-intent" debate at Thumbtack, when we rolled out changes that A/B tests showed increased conversion (by a lot), but some people in the company thought the changes were ugly and "off-brand" and argued they brought in the wrong type of customers. So we ran the test again that we knew raised conversion by a lot, and then followed th…

> Let's people check test too often and allows them to stop the test when it hits a good result I admit to attempting to be guilty of this in the past and being stopped by our analytics team (in the sense that they took the time to patiently explain to me why what I was doing was statistically unsound). It's not obvious, IMO.

Yeah, classical A/B testing only works if you specify the duration of time that the experiment will run in advance and stick to it. There are ways that you can continuously monitor how an experiment is performing (Google "sequential Bayesian analysis") but the math is quite a bit more complex and nuanced.

Honestly, for most startups a simple multi-armed bandit approach is probably the way to go. Don't worry about statistical significance; just throw some "lite" reinforcement learning on top of your product's aesthetics and enjoy the incremental profit. (Caveat: do not apply MAB to major product changes.)

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#52
post #5

Earlier quoted context omitted.

> Let's people check test too often and allows them to stop the test when it hits a good result I admit to attempting to be guilty of this in the past and being stopped by our analytics team (in the sense that they took the time to patiently explain to me why what I was doing was statistically unsound). It's not obvious, IMO.

Yeah, classical A/B testing only works if you specify the duration of time that the experiment will run in advance and stick to it. There are ways that you can continuously monitor how an experiment is performing (Google "sequential Bayesian analysis") but the math is quite a bit more complex and nuanced. Honestly, for most startups a simple multi-armed bandit approach is probably the way to go. Don't worry about sta…

> Honestly, for most startups a simple multi-armed bandit approach is probably the way to go.

Doesn't this mean maintaining all the variants forever?

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#53

A/B tests tell you about short term gains, but don't tell you about long term issues you may be accumulating due to things like dark patterns, clickbait headlines, shoddy article topics and more. A/B tests don't take into account the loss of prestige or reputation that the options give. I've seen this repeatedly with ArsTechnica, which has devolved into so much political and clickbait material that I don't even reall…

That's so true!

Too many companies just optimize for short term instead of long term.

Yet, on the other hand it's completely natural: Not every founder wants to stay at their company for ever. Not all companies have the objective of staying alive. It's much more exciting and thrilling to chase that hockey stick growth and IPO/buyout.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#55
post #5
post #3

I could not disagree with this more. I remember vividly having this "low-intent" vs "high-intent" debate at Thumbtack, when we rolled out changes that A/B tests showed increased conversion (by a lot), but some people in the company thought the changes were ugly and "off-brand" and argued they brought in the wrong type of customers. So we ran the test again that we knew raised conversion by a lot, and then followed th…

> Let's people check test too often and allows them to stop the test when it hits a good result I admit to attempting to be guilty of this in the past and being stopped by our analytics team (in the sense that they took the time to patiently explain to me why what I was doing was statistically unsound). It's not obvious, IMO.

Bless you for listening. I find statistics to be one of those subjects that can be very counter intuitive and hard to grok, which leads to people just shutting down when the topic comes up. I have a hard enough time convincing people that just the use of statistical average is inferior to median in nearly every context they care about in business, let alone why stopping a test early invalidates it.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#56
post #40
post #29

Earlier quoted context omitted.

Meh. There is statistics for the purpose of uncovering Truth, and statistics for the purpose of making a business decision. The difference is that when we talk about Truth, a small error is still an error. When we make business decisions, it is fine to make a decision that is probably right, and we know isn't far wrong. Here is a perfectly valid test procedure that illustrates the difference. Decide the most time you…

Truth isn't that different from business. Both are statistical. Making a decision that is not statistically valid is bad business, as you could be making things worth as often as better.

I think you might be missing btilly's point. He is saying that the testing protocol should factor in the cost of being wrong. These costs might be drastically different for a business and for a scientific pursuit.

If A is ahead of B by a hair and the my flawed protocol chooses B the cost to business might not be that high. But the same protocol might not be a good one if the cost of making that mistake is very high. The probabilities of the errors remain the same for the two scenarios, the expected costs are different.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#57
post #29

Earlier quoted context omitted.

Meh. There is statistics for the purpose of uncovering Truth, and statistics for the purpose of making a business decision. The difference is that when we talk about Truth, a small error is still an error. When we make business decisions, it is fine to make a decision that is probably right, and we know isn't far wrong. Here is a perfectly valid test procedure that illustrates the difference. Decide the most time you…

I mean I think the only reason it would horrify your analytics team is that if you want to do something that sophisticated, you may as well just use a better multi-armed bandit function like Thompson Sampling or UCB-1 (which is very, very similar to what you've described, although more formalized). So I think its wrong to say that you'd never find it in a stats class.

Bandits and A/B are meant for solving very different problem.

In bandit there is a clear explore/exploit trade off. There is no such trade off in the A/B formulation, although it does get used in scenarios that have such trade offs.

If I can pull the lever a finite and small number of times there is a strong incentive for using a bandit. In this case I don't want to pull the wrong lever as few times as possible. On the other hand if I am given an unlimited number of pulls, I can afford to pull the wrong one many times more (still finite) for the sake of 'knowledge' knowing well that I would have infinitely many opportunities to exploit that knowledge.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#58
post #57

Earlier quoted context omitted.

I mean I think the only reason it would horrify your analytics team is that if you want to do something that sophisticated, you may as well just use a better multi-armed bandit function like Thompson Sampling or UCB-1 (which is very, very similar to what you've described, although more formalized). So I think its wrong to say that you'd never find it in a stats class.

Bandits and A/B are meant for solving very different problem. In bandit there is a clear explore/exploit trade off. There is no such trade off in the A/B formulation, although it does get used in scenarios that have such trade offs. If I can pull the lever a finite and small number of times there is a strong incentive for using a bandit. In this case I don't want to pull the wrong lever as few times as possible. On t…

And for a business there always is. In an a/b test that is to improve customer conversion, in a perfect world, you use the superior method on everyone, converting the maximum number of customers. That saves you money.

In other words, the opportunity cost of putting someone in the wrong group creates such a trade-off. You can pull the lever as many times as you want, but each one potentially costs you money. It's textbook bandits.

Re: Conservation of Intent: why A/B tests aren’t as effective as they look

#59
post #57

Earlier quoted context omitted.

Bandits and A/B are meant for solving very different problem. In bandit there is a clear explore/exploit trade off. There is no such trade off in the A/B formulation, although it does get used in scenarios that have such trade offs. If I can pull the lever a finite and small number of times there is a strong incentive for using a bandit. In this case I don't want to pull the wrong lever as few times as possible. On t…

And for a business there always is. In an a/b test that is to improve customer conversion, in a perfect world, you use the superior method on everyone, converting the maximum number of customers. That saves you money. In other words, the opportunity cost of putting someone in the wrong group creates such a trade-off. You can pull the lever as many times as you want, but each one potentially costs you money. It's text…

To a large extent I agree with you.

Differences creep in when there is ambiguity and judgement involved on what is that metric that the org wants to optimize. This is fairly common. Typically, in these situations its the PMs who make the final call. There the goal of the experiment protocol is glean as much knowledge as possible, and present it to the PM. The thinking there is -- if that comes at the cost of exposing some customers to bad choices, so be it.

Post reply on HN