Earlier quoted context omitted.
Why would this not be a good domain for image recognition?
If you’re trying to classify images as whether or not they contain something, you need a bunch of images that you already know don’t contain the thing (got it), and a bunch of images that you already know /do/ contain the thing. We lack the latter, so there’s nothing for the algorithm to learn based on.
In fact, this idea seems so obvious that it must have been tried already…