A few comments below talk about how tokenizing images using stuff like CLIP de-facto yields blurry image descriptions, and so these are ‘blind’ by some definitions. Another angle of blurring not much discussed is that the images are rescaled down; different resolutions for different models. I wouldn’t be surprised if Sonnet 3.5 had a higher res base image it feeds in to the model. Either way, I would guess that we’ll…
It clearly wasn’t trained on this task and suffers accordingly.
However, with chatgpt, it will create python to do the analysis and has better results.