Thanks for the comment :D
> There is no “testing” going on here, unfortunately. You’re just taking turns placing users in the treatment or control group, arbitrarily, depending on when they visit the page. How can you measure any treatment effect when everyone is part of either group?
This is not a perfect solution, but I have found that it is good enough to be able to identify clear winners (if there is actually a winning version). A lot of followers won't visit your profile multiple times anyway. They visit it, and they either follow it or they don't - and they will of course be influenced by the currently displayed version :). They won't just come back to your profile over and over again for no reason, but if they do, they will also convert better on the version they prefer. So a version might nudge the user into following you while another one might not. I would disagree that this is not testing. I think it is for the majority of the profile clicks you receive.
> I guess you could try something like switchback testing, but I’m not convinced visits to the average Twitter profile will yield enough samples.
I don't think that would be possible with the current Twitter API capabilities anyway.
> I think it’s a well-executed idea, but I don’t think it’s fair to sell results under the guise of statistical validity when they don’t appear to have that. (although it’s just Twitter profiles, not eg medical treatment, so no real harm done)
While the results might not be perfectly accurate, I think they are accurate enough to provide value, especially if you let the test run long enough to get a big sample size. I personally use Birdy (obviously :D) and I have noticed much better conversion, which is why I'm confident.
I'm looking forward to seeing new capabilities appear on the Twitter API to always make the process more accurate though.