>
it's quite possible that A/B testing has a long-term effect of pushing those participants whose tolerance has been exceeded out of the study population entirely.I'd compare this to how evaporation rate increases with temperature, as more particles find themselves with enough energy to escape the liquid.
From my personal experience, even if I can tolerate a lot of UX abuse, each such "optimization" lowers my threshold of switching to a competitor. Software in general, and SaaS specifically, resists commoditization, but every now and then an actual alternative to a product/service I'm using shows up - and whether or not I switch (and when) is correlated with how much I resent the incumbent for their UX "improvements".
I'd add one bullet point to your list:
- Unlike regular scientific experimentation, A/B testing is a methodology primarily spread in business circles using regular hype channels. That is, the average practitioner is not qualified to execute it correctly, which is one of the reasons I see A/B testing more as tools to launder arbitrary decisions. Because consequences of doing it wrong are typically not immediately apparent or obvious, both companies and customers suffer (and a vast space for fraudsters is created).
I'm in a charitable mood, so I'm not passing judgement on people for not having PhD-level understanding of statistics - just pointing out that, to the degree much larger than in sciences (even soft ones, which suffer some of the same structural problems), there's little pressure to do such tests correctly (and there's lot of ways to make money or status by doing them without regards for correctness).
From what I hear, a common way of executing A/B test badly and getting bullshit results, is by terminating the test early when it shows the relevant metrics improving for the test group - vs. running it longer if no big improvements are observed (or the metrics start getting worse for the test group). This biases the experiment towards giving false positives. This problem was big enough that there was a debacle around Optimizely few years ago, whose UI was accused to promote this early termination of tests. The cynical take I'm still somewhat partial to is that it wasn't an accident (if not done on purpose, then possibly... a result of an A/B test!) - false positives make the (statistically naive) users feel they're getting more value from Optimizely than they actually are.