This article would not come as a surprise to anyone who works with ConvNets. Sadly, that might not the case for those outside of the field, largely due to media's inadequate coverage of our advances (but this is common outside our field too). No one in the field really believes ConvNets see better than humans. They are very good single glance texture recognizers. It's as if you flashed an image and looked at it for a split second without giving yourself a chance to look around and take some time to gain any higher-level scene understanding. If you tried this with this image you might also think you had seen a leopard. Another point to make is not from modeling side but from data side. If in the training data the leopard texture is highly indicative of leopard, then the ConvNet will learn to strongly associate it as such. As the article mentions, a quick hack would be to make sure that your training data contains many leopard-textured items of different classes. You might then expect the ConvNet to seek other features to latch on to and become less reliant on the texture itself.
Also, we carried out an experiment on ImageNet and the outcome was that "One human labeler (me, incidentally) with a fixed amount of training and a slightly-above average determination reached ~5% top-5 error on a subset of ImageNet test set". The media sees this and it immediately gets spun to "AI now Super-Human. And we're all going to die." It makes a lot of us cringe every time.
Many people in Computer Vision now consider ImageNet "squeezed" out of juice - we're good at texture recognition and when an object is in plain view, and we're now searching for harder tasks and more dynamic range with respect to human performance, in areas such as harder 3D/Spatial tasks, Image Captioning, Visual Q&A, etc. The hope is that these harder datasets might in turn guide us in developing models with more nuanced understanding.