Earlier quoted context omitted.
The captcha example isn't exactly fair, the images shown can be newly generatrd, repeated, or used as control images, and even if they are unclassifiable then you're not seeing the successes. Without knowing the whole system there's too much hidden bias to claim the computer is more of less accurate than a commuter.
No doubt all of your excuses are at least partially true, however, at this point, literally after almost two decades of billions of people training these things to see bikes, it still needs a lot more work.
In fact, visual challenges like this are a mostly solved problem, and the real magic is the classification that happens behind the scenes.