Well, as a blind user, I'd like to point at the OpenAI Vision integration of BeMyEyes! Being able to get fully detailed scene descriptions including OCR and translation all in one package was pretty much a game changer for me. Not so much kick-ass, but still works nicely: https://github.com/mlang/tracktales -- My MPD track announcer with support for describing album art...
Can I ask if you think this might make alt text on images obsolete? Do you use the alt text where it’s available, or BeMyEyes (I presume you have a choice)?
While alt texts could theoretically be replaced by a browser/screen reader functionality that asks a vision model to describe the image, it is a waste of time and energy to have each and every user do it over and over again.