1. If the effect size is incredibly small, which most minor UI changes will be, finding statistical tests to prove them is really difficult. If you’re looking for an incredibly small positive effect, even with hundreds of thousands of sample size, the probability of rejecting the null hypothesis while the effect is actually a negative influence is surprisingly high! Very easy to make mistakes.
2) short term gains on engagement may lead to long term disengagement.
3) business incentives for management are easily misaligned. I would imagine a dominant negative influence is managers exaggerating the statistical influence found in a test because that means they get to lead the change, on an otherwise vast tech ecosystem the performance of which probably won’t change all that much. Attribution is also hard (how sure are they on how much to attribute here?) so credit is difficult to allocate beyond initial value sizing.