Earlier quoted context omitted.
The null hypothesis is that the participants are just as able to find those cat-like features in either perturbation of the image, and would pick the “right” one only 50% of the time.
No, the null hypothesis is that neither image elicits a cat-like response. It requires asking an open-ended question such as "what does this image look like to you?" Once you prime the subject, you have artificially restricted the responses to "Not really, sure, kinda?" Remember, the ML model is objectively selecting cat with (very) high probability out of the entire corpus of possible responses. The human should be…
Because the answer to that is simply “it looks like flovers in a vase”. There is no question about human’s ability to tell what the image is.
So much so that if you ask the humans to describe the images they would probably say something along the lines of “two identical images of the same flowers”.
So you would think if you ask them which one is more cat like they will shrug and pick one at random. Since it is a nonsense question. Yet people were able to pick up the manipulated image as more cat-like. Which means there is some signal they are able to pick up on.
> it suggests a fairly large gap between human and machine perception
Naturally. That is not at dispute, neither is it the subject of this study.