Earlier quoted context omitted.
> I guess they do "see" but more like "see an explanation of the image", not "see" as in experience visually. Images are tokenized and fed to the exact same model, they can “visually inspect” images, eg “find the 2 differences between two images” and “where’s Waldo”-style things. So your mental model that they see descriptions is inaccurate.
> Images are tokenized Exactly, here is where the fidelity of an image is being lost, they don't "see" visually, they get a representation of the image via tokens, that's why I said they don't see but basically "see an explanation of the image". I don't mean like a caption, but in the end, they act and work with tokens, not pixels or actual images, internally. Example from Grok and Claude, with a very simple test cas…
Re: Tell HN: Dont use Claude Design, lost access to my projects after unsubscribing
#101You don’t “see” visually either! It’s just that when photons hit your rods and cones some electrical impulses go down your optic nerves and hit your visual cortex, and some math happens that your sensory systems interpret as vision. But it’s nothing of the sort, just a low-fidelity trick.