I could not disagree with this more. I remember vividly having this "low-intent" vs "high-intent" debate at Thumbtack, when we rolled out changes that A/B tests showed increased conversion (by a lot), but some people in the company thought the changes were ugly and "off-brand" and argued they brought in the wrong type of customers. So we ran the test again that we knew raised conversion by a lot, and then followed th…
> Let's people check test too often and allows them to stop the test when it hits a good result I admit to attempting to be guilty of this in the past and being stopped by our analytics team (in the sense that they took the time to patiently explain to me why what I was doing was statistically unsound). It's not obvious, IMO.
Conservation of Intent: why A/B tests aren’t as effective as they look
11–20 of 59 posts
Re: Conservation of Intent: why A/B tests aren’t as effective as they look
#12I could not disagree with this more. I remember vividly having this "low-intent" vs "high-intent" debate at Thumbtack, when we rolled out changes that A/B tests showed increased conversion (by a lot), but some people in the company thought the changes were ugly and "off-brand" and argued they brought in the wrong type of customers. So we ran the test again that we knew raised conversion by a lot, and then followed th…
Re: Conservation of Intent: why A/B tests aren’t as effective as they look
#13I could not disagree with this more. I remember vividly having this "low-intent" vs "high-intent" debate at Thumbtack, when we rolled out changes that A/B tests showed increased conversion (by a lot), but some people in the company thought the changes were ugly and "off-brand" and argued they brought in the wrong type of customers. So we ran the test again that we knew raised conversion by a lot, and then followed th…
> Let's people check test too often and allows them to stop the test when it hits a good result I admit to attempting to be guilty of this in the past and being stopped by our analytics team (in the sense that they took the time to patiently explain to me why what I was doing was statistically unsound). It's not obvious, IMO.
If you create a culture where positive A/B tests are lauded (which is good!), then you create a lot of people who want A/B tests to finish in positive ways. For those people, it doesn't really matter if their A/B test actually improves things, only if it looks like it does. This isn't nefarious, this is just human nature, but it creates a lot of creativity and energy at finding ways to making winning A/B tests. We'd have people run a test, where their new variation got off to a bad start, then say "oh, it was a bug", then they'd fix some irrelevant thing and start over just to reset the counters. I was guilty of it sometimes. You get excited about your tests and want them to win. That is why it is critical either to have gatekeepers like an analytics team to keep you honest or have a really specific protocol on how your company runs tests and only consider results of tests that followed the protocol.
Re: Conservation of Intent: why A/B tests aren’t as effective as they look
#14Slightly relevant but useful: use meditation modeling ( https://eng.uber.com/mediation-modeling/ )
Re: Conservation of Intent: why A/B tests aren’t as effective as they look
#15The reason economists talk about opportunity cost is because people are constantly optimizing decisions based on new information. (Humans may not deal with prices and numbers very well, but they're pretty well evolved to break time into chunks and work out plans to solve problems.)
If you talk to an individual, they might say "I can't afford it," or you may talk to someone who didn't click through and they might say, "I was just browsing." The fallacy behind both is you're creating archetypes and assuming they represent the modes of the population.
And even if you talk to the individuals you based those archetypes on, there is a whole history behind how they arrived at "I can't afford it." Those changing circumstances are why the aggregate behavior doesn't show some arbitrary level of "affordability," and instead you see a smooth curve of consumer demand.
And the opportunity cost of continuing to view a web page will not have neatly quantized levels of intent, but rather individuals have a broad array of competing interests.
Re: Conservation of Intent: why A/B tests aren’t as effective as they look
#16Earlier quoted context omitted.
....and criticising statistics is the laziest kind of scientific criticism.... A/B tests are fine. They work. They allow inference of causality. They are easy to understand, and can be fun to run. They get you 90% of wherever you want to go, and such over-the-top criticism just seems like badly executed pretentiousness.
You’re partly right - a scientific endeavor to figure out the color of a button would be over-the-top, because it’s not that important. But to the article’s point, if you’re running banking software or something, your users don’t give a shit what the button colours are; they will slog through whatever you develop because they need to get stuff done. A/B tests are a small tool that sometimes get taken too far or used…
In the language of this article, banking app users are mostly all 'high intent'. But that doesn't mean you can't evaluate criteria other than users who completed the workflow to determine what is a design improvement. You can still measure time to completion, how long it took the user from entering the workflow to completing a given task as a measure, and play with the design. Optimize the things users are doing the most, that sort of thing. A/B testing can help you there. It's not the be all end all; you still need sound UX design to figure out what designs to test out, but it can give you measurable data as to what works, rather than just UX gut feeling, or purely lab based results which don't reflect reality.
Re: Conservation of Intent: why A/B tests aren’t as effective as they look
#17I've seen this repeatedly with ArsTechnica, which has devolved into so much political and clickbait material that I don't even really visit anymore. Yes, I'm guity myself of clicking on those articles when I do visit, but at a certain point I've found that Ars doesn't have the news I'm after, so I turn elsewhere and now Ars has one less viewer.
Re: Conservation of Intent: why A/B tests aren’t as effective as they look
#18Let's say you're actually trying to optimize total transaction value on the site or total number of transactions or something like the overall fraction of users with at least one transaction within a certain window of time. Then - as the article rightly observes - getting users not to bounce on a particular page is a TERRIBLE proxy to what you're optimizing for. If that's not clear to you, you have no business running A/B tests without supervision.
Source: co-designed one iteration of the experimentation framework for Booking.com many years ago. Indirectly managed the team of much more qualified people that took it a world further.
Re: Conservation of Intent: why A/B tests aren’t as effective as they look
#19I could not disagree with this more. I remember vividly having this "low-intent" vs "high-intent" debate at Thumbtack, when we rolled out changes that A/B tests showed increased conversion (by a lot), but some people in the company thought the changes were ugly and "off-brand" and argued they brought in the wrong type of customers. So we ran the test again that we knew raised conversion by a lot, and then followed th…
Attaching A/B test variations, sourcing data (UTMs), and other data to actual customer/user records in your database is a great start to understanding your results better. That allows you to do things like suss out "intent" (essentially conversion rate), "quality" (essentially LTV), etc. down the road and really understand the value of your tests and acquisition channels.
I'm always surprised when companies with more than enough resources to do this don't do this.
Re: Conservation of Intent: why A/B tests aren’t as effective as they look
#20A/B tests tell you about short term gains, but don't tell you about long term issues you may be accumulating due to things like dark patterns, clickbait headlines, shoddy article topics and more. A/B tests don't take into account the loss of prestige or reputation that the options give. I've seen this repeatedly with ArsTechnica, which has devolved into so much political and clickbait material that I don't even reall…
Yet I'd like to add that I don't think that testing frameworks (at this point it would be misleading to call them strictly A/B testing) HAVE to only reflect short term gains. It's hard to come up with (proxy) metrics that hold that kind of short term optimization in check.
Just to pick one fairly obvious example for e-commerce: you can track returns/customer service contacts as a health metric or even explicitly assign a value to them to fold them into the primary metric you're optimizing.
These kinds of safeguards typically need longer recording periods, so it's potentially a lot of data science and engineering effort to build tools that can handle the long term data collection and analysis. But it's not impossible. It's rather something people love to pretend is not a problem.