This actually doesn't seem like it's a giant lift using modern image classifiers. The basic idea is to use diffusion classifiers to caption the image to generate descriptive text and append the prompt. The work part is getting the ensemble right since you'll need to use a general classifier, like BLIP, to identify say a bunch of text from a plant and then, in this example, use structured OCR and pl@ntnet to get more…
It's like the difference between you telling me what's outside the window and then asking me questions about it – versus me being able to look out the window myself.