Earlier quoted context omitted.
I'll take a crack at an explanation: Google has machine learning technology that automatically reads street address numbers on houses, regardless of angle, color, size, focus, or typeface. Using the same type of algorithm to answer "does this ad have anything that looks like a button in it?" is relatively simple by at least two orders of magnitude. So rather than it being "easy" in some absolute sense, what's meant h…
Lots of things have button and aren't maliciously masquerading as download links. Furthermore, if they start doing that then ad producers will start changing their ad images to evade the algorithm. It's much harder to use ML for a problem when you have a malicious opponent actively working against it.
And changing the images to evade the algorithm would also directly affect how likely they are to trick people, so it's win/win?