This seems ridiculous. Based on the output of an automated tool developed by 3M that produces a saliency map based on an image, I'm not sure how we can conclude that people ignore skyscrapers.
How does the software work? Seam carving style "energy functions" or neural networks? If it's neural networks how did they train it? Eye tracker?
https://web.archive.org/web/20180604223618/http://solutions....
Validation data is eye tracking
http://multimedia.3m.com/mws/media/1006827O/3msm-visual-atte...