First, you really should move away from frequentist statistical testing and use Bayesian statistics instead. It is perfect for such occasions where you want to adjust your beliefs in what UX is best based on empirical data to support your decision. With collecting data you are increasing confidence in your decision rather than trying to meet an arbitrary criterion of a specific p-value. Second, the “run-in-parallel”…
The frequentist/bayesian debate is not one I understand well enough to opine - do you have any reading you'd recommend for this topic?