Live data from Hacker News

Vision language models are blind

vlmsareblind.github.io

171–180 of 202 posts

Re: Vision language models are blind

#171

Entertaining, but I think the conclusion is way off. > their vision is, at best, like that of a person with myopia seeing fine details as blurry is a crazy thing to write in an abstract. Did they try to probe that hypothesis at all? I could (well actually I can't) share some examples from my job of GPT-4v doing some pretty difficult fine-grained visual tasks that invalidate this. Personally, I rate this paper [1], wh…

> these huge GenAI models are pretty good at things

Is this the sales pitch though? Because 15 years ago, I had a scanner with an app that can scan a text document and produce the text on Windows. The machine had something like 256Mb of RAM.

Tech can be extremely good at niches in isolation. You can have an OCR system 10 years ago and it'll be extremely reliable at the single task it's configured to do.

AI is supposed to bring a new paradigm, where the tech is not limited to the specific niche the developers have scoped it to. However, if it reliably fails to detect simple things a regular person should not get wrong, then the whole value proposition is kicked out of the window.

Re: Vision language models are blind

#172

Earlier quoted context omitted.

Exactly... I've found GPT-4o to be good at OCR for instance... doesn't seem "blind" to me.

I know that GPT-4o is fairly poor to recognize music sheets and notes. Totally off the marks, more often than not, even the first note is not recognize on a first week solfège book. So unless I missed something but as far as I am concerned, they are optimized for benchmarks. So while I enjoy gen AI, image-to-text is highly subpart.

Most adults with 20/20 vision will also fail to recognize the first note on a first week solfege book.

Re: Vision language models are blind

#174
post #54

Earlier quoted context omitted.

Yea, really if you look at human learning/seeing/acting there is a feedback loop that LLM for example isn't able to complete and train on. You see an object. First you have to learn how to control all your body functions to move toward it and grasp it. This teaches you about the 3 dimensional world and things like gravity. You may not know the terms, but it is baked in your learning model. After you get an object you…

Embodied is useful, but I think not necessary even if you need learning in a 3D environment. Synthesized embodiment should be enough. While in some cases[0] it may have problems with fidelity, simulating embodied experience in silico scales much better, and more importantly, we have control over time flow . Humans always learn in real-time, while with simulated embodiment, we could cram years of subjective-time exper…

> Humans always learn in real-time

In the sense that we can't fast-forward our offline training, sure, but humans certainly "go away and think about it" after gaining IRL experience. This process seems to involve both consciously and subconsciously training on this data. People often consciously think about recent experiences, run through imagined scenarios to simulate the outcomes, plan approaches for next time etc. and even if they don't, they'll often perform better at a task after a break than they did at the start of the break. If this process of replaying experiences and simulating variants of them isn't "controlling the flow of (simulated) time" I don't know what else you'd call it.

Re: Vision language models are blind

#175
post #45

Speaking as someone with only a tenuous grasp of how VLMs work, this naïvely feels like a place where the "embodiement" folks might have a point: Humans have the ability to "refine" their perception of an image iteratively, focusing in on areas of interest, while VLMs have to process the entire image at the same level of fidelity. I'm curious if there'd be a way to emulate this (have the visual tokens be low fidelity…

Humans are actually born with blurry vision as the eye takes time to develop, so human learning starts with low resolution images. There is a theory that this is not really a limitation but a benefit in developing our visual processing systems. People in poorer countries that get cataracts removed when they are a bit older and should at that point hardware-wise have perfect vision do still seem to have some lifelong deficits.

It's not entirely known how much early learning in low resolution makes a difference in humans, and obviously that could also relate more to our specific neurobiology than a general truth about learning in connectionist systems. But I found it to be an interesting idea that maybe certain outcomes with ANNs could be influenced a lot by training paradigms s.t. not all shortcomings could be addressed with only updates to the core architecture.

Re: Vision language models are blind

#176

Entertaining, but I think the conclusion is way off. > their vision is, at best, like that of a person with myopia seeing fine details as blurry is a crazy thing to write in an abstract. Did they try to probe that hypothesis at all? I could (well actually I can't) share some examples from my job of GPT-4v doing some pretty difficult fine-grained visual tasks that invalidate this. Personally, I rate this paper [1], wh…

There's definitely something interesting to be learned from the examples here - it's valuable work in that sense - but "VLMs are blind" isn't it. That's just clickbait.

Re: Vision language models are blind

#177

Entertaining, but I think the conclusion is way off. > their vision is, at best, like that of a person with myopia seeing fine details as blurry is a crazy thing to write in an abstract. Did they try to probe that hypothesis at all? I could (well actually I can't) share some examples from my job of GPT-4v doing some pretty difficult fine-grained visual tasks that invalidate this. Personally, I rate this paper [1], wh…

There are quite a few "ai apologists" in the comments but I think the title is fair when these models are marketed towards low vision people ("Be my eyes" https://www.youtube.com/watch?v=Zq710AKC1gg ) as the equivalent to human vision. These models are implied to be human level equivalents when they are not. This paper demonstrates that there are still some major gaps where simple problems confound the models in unex…

I don’t see Be My Eyes or other similar efforts as “implied” to be equivalent to humans at all. They’re just new tools which can be very useful for some people.

“These new tools aren’t perfect” is the dog bites man story of technology. It’s certainly true, but it’s no different than GPS (“family drives car off cliff because GPS said to”).

Re: Vision language models are blind

#178

Earlier quoted context omitted.

It really shouldn't be surprising that these models fail to do anything that _they weren't trained to do_. It's trivially easy to train a model to count stuff. The wild thing about transformer based models is that their capabilities are _way_ beyond what you'd expect from token prediction. Figuring out what their limitations actually are is interesting because nobody fully knows what their limitations are.

I agree that these open ended transformers are way more interesting and impressive than a purpose built count-the-polygons model, but if the model doesn't generalize well enough to figure out how to count the polygons, I can't be convinced that they'll perform usefully on a more sophisticated task. I agree this research is really interesting, but I didn't have an a priori expectation of what token prediction could ac…

> I agree that these open ended transformers are way more interesting and impressive than a purpose built count-the-polygons model, but if the model doesn't generalize well enough to figure out how to count the polygons, I can't be convinced that they'll perform usefully on a more sophisticated task.

I think people get really wrapped into the idea that a single model needs to be able to do all the things, and LLMs can do a _lot_, but there doesn't actually need to be a _one model to rule them all_. If VLMs are kind of okay at image intepretation but not great at details, we can supplement them with something that _can_ handle the details.

Re: Vision language models are blind

#180

This is an interesting article and goes along with how I understand how such models interpret input data. I'm not sure I would characterize the results as blurry vision, but maybe an inability to process what they see in a concrete manner. All the LLMs and multi-modal models I've seen lack concrete reasoning. For instance, ask ChatGPT to perform 2 tasks, to summarize a chunk of text and to count how many words are in…

I hope you are aware of the fact that LLMs does not have direct access to the stream of words/characters. It is one of the most basic things to know about their implementation.

Yes, but it could learn to associate tokens with word counts as it could with meanings.

Even still, if you ask it for token count it would still fail. My point is that it can’t count, the circuitry required to do so seems absent in these models

Post reply on HN