Hypothetically, could this be fixed by changing the input method. For instance, I just quickly looked up how humans process imagery. "the primary visual cortex, located at the back of the brain, receives the visual signals and processes basic visual features like edges, lines, and orientations." So, potentially if we did a pre-processing step to get more features out beforehand we would see different results in the o…
The way to fix this is simpler: ensure counter-factuals are present in the training data, then the VLM will learn not to be dependent on its language priors/knowledge.