The net correctly identified "leopard". Was it taught about sofas? Who knows, maybe Sofa had a high score as well on the output.
Or, look at the Dalmatian/Cherry picture. The net identified "Dalmatian" which is a 100% valid response! But whoever labeled it wanted "cherry". The picture is 50% cherry 50% dalmatian.
Pictures often have more than one element and a pure ConvNet is "one picture to one label"