Earlier quoted context omitted.
I'm not really sure what you are commenting on. They tested this across multiple pairs of categories: (sheep, chair), (dog, bottle), (cat, truck), and (elephant, clock). This isn't a phenomena related to cats. The whole point of the study is to measure the impact of the noise. The "baseline" or control here would be to to not add noise to either of the two images and arbitrarily label one "cat" and the other "truck"…
1) It is not adding pure noise. 2) If humans when prompted tend to always see something more in one picture than the other when random noise is added, the baseline might not be 50/50 as no matter what you ask you get a systematic preference. Double blinding would not remove this.
I understand that the concern could be something like "random perturbations of cat images make every image simultaneously less cat-like and hence more like anything else". My opinion is that Experiment 4 (making an image more cat-like or truck-like) covers concerns of this nature. Even if there are two random perturbations where one makes an image cat-like and the other more truck-like, it is completely arbitrary whether you label the perturbation as cat-like or truck-like (since they are randomly generated). That means, even if you measure a difference (and even if two random perturbations have larger differences than this construction!), you cannot control the direction. This method gives you a way to control it. Personally, I don't think this study is about measuring the influence over some baseline. It's about showing that you can indeed choose the direction of influence.