To be clear, it depends on what your definition of "insanely bad" is.
I'd say ChatGPT/Gemini make egregious mistakes on ~10% of my photo uploads.
I recently uploaded a photo of a short-billed dowitcher and ChatGPT told me that it was a Wilson's snipe, explaining all sorts of details about the legs and tail feathers (neither of which were visible in my pic!).
I then followed up explaining that a Wilson's snipe hadn't been seen at my location since last November (and that Wilson's snipe was out of season at my location) and Chat revised its estimate downward to 85% Wilson's snipe.
Again, I followed up and I revealed the precise location of the bird and ChatGPT said something like "oh yeah, 99% short-billed dowitcher"!
I've had similar experiences w/ Gemini (haven't tested Claude).
Again, 90% success rate is pretty good, but the other 10% of the time, the 2 LLMs that I use fail on species ID and often hallucinate features on bird photos.
edits for typos, plus another example from the same "birding outing" the other day.
I uploaded a very clear photo of a sparrow.
* ChatGPT says "song sparrow"
* I explain, "no way. this sparrow has yellow over its eye and the breast is wrong for song sparrow."
* ChatGPT: Oh yeah, savannah sparrow
* I explain, beak is too big for savannah sparrow.
* ChatGPT: Oh yeah, saltmarsh sparrow.
* I expalin, "no orange on the bird's face."
* ChatGPT: oh yeah, seaside sparrow (finally correct!)