Live data from Hacker News

Vision language models are blind

vlmsareblind.github.io

101–110 of 202 posts

Re: Vision language models are blind

#101

Earlier quoted context omitted.

There are quite a few "ai apologists" in the comments but I think the title is fair when these models are marketed towards low vision people ("Be my eyes" https://www.youtube.com/watch?v=Zq710AKC1gg ) as the equivalent to human vision. These models are implied to be human level equivalents when they are not. This paper demonstrates that there are still some major gaps where simple problems confound the models in unex…

Ah yes the blind person who constantly needs to know if two lines intersect. Let's just ignore what a blind person normally needs to know. You know what blind people ask? Sometimes there daily routine is broken because there is some type of construction and models can tell you this. Sometimes they need to read a basic sign and models can do this. Those models help people already and they will continue to get better.…

As an aside... from 2016 this is what was a valid use case for a blind person with an app.

Seeing AI 2016 Prototype - A Microsoft research project - https://youtu.be/R2mC-NUAmMk

https://www.seeingai.com are the actual working apps.

The version from 2016 I recall showing (pun not intended) to a coworker who had some significant vision impairments and he was really excited about what it could do back then.

---

I still remain quite impressed with its ability to parse the picture and likely reason behind it https://imgur.com/a/JZBTk2t

Re: Vision language models are blind

#102
post #9
post #7

Wow, that is embarrassingly bad performance for current SOTA models (GPT-4o, Gemini-1.5 Pro, Sonnet-3, Sonnet-3.5), which are advertised and sold as being able to understand images, e.g., for guiding the blind or tutoring children in geometry! The tasks at which they fail are ridiculously simple for human beings, including, for example: * counting the number of times two lines intersect; * detecting whether two circl…

I don't see how this is "embarrassing" in the slightest. These models are not human brains, and the fact that people equate them with human brains is an embarrassing failure of the humans more than anything about the models. It's entirely unsurprising that there are numerous cases that these models can't handle that are "obvious to humans." Machine learning has had this property since its invention and it's a classic…

> is an embarrassing failure of the humans more than anything about the models

No, it's a failure of the companies who are advertising them as capable of doing something which they are not (assisting people with low vision)

Re: Vision language models are blind

#104
post #9

Earlier quoted context omitted.

I don't see how this is "embarrassing" in the slightest. These models are not human brains, and the fact that people equate them with human brains is an embarrassing failure of the humans more than anything about the models. It's entirely unsurprising that there are numerous cases that these models can't handle that are "obvious to humans." Machine learning has had this property since its invention and it's a classic…

> is an embarrassing failure of the humans more than anything about the models No, it's a failure of the companies who are advertising them as capable of doing something which they are not (assisting people with low vision)

But they CAN assist people with low vision. I've talked to someone who's been using a product based on GPT-4o and absolutely loves it.

Low vision users understand the limitations of accessibility technology better than anyone else. They will VERY quickly figure out what this tech can be used for effectively and what it can't.

Re: Vision language models are blind

#105

My guess is that the systems are running image recognition models, and maybe OCR on images, and then just piping that data as tokens into an LLM. So you are only ever going to get results as good as existing images models with the results filtered through an LLM. To me, this is only interesting if compared with results of image recognition models that can already answer these types of questions (if they even exist, I…

That's not how they work. The original GPT-4 paper has some detail: https://cdn.openai.com/papers/gpt-4.pdf

Or read up on PaliGemma: https://github.com/google-research/big_vision/blob/main/big_...

Re: Vision language models are blind

#106

Entertaining, but I think the conclusion is way off. > their vision is, at best, like that of a person with myopia seeing fine details as blurry is a crazy thing to write in an abstract. Did they try to probe that hypothesis at all? I could (well actually I can't) share some examples from my job of GPT-4v doing some pretty difficult fine-grained visual tasks that invalidate this. Personally, I rate this paper [1], wh…

There are quite a few "ai apologists" in the comments but I think the title is fair when these models are marketed towards low vision people ("Be my eyes" https://www.youtube.com/watch?v=Zq710AKC1gg ) as the equivalent to human vision. These models are implied to be human level equivalents when they are not. This paper demonstrates that there are still some major gaps where simple problems confound the models in unex…

I disagree. I think the title, abstract, and conclusion not only misrepresents the state of the models but it misrepresents Thier own findings.

They have identified a class of problems that the models perform poorly at and have given a good description of the failure. They portray this as a representative example of the behaviour in general. This has not been shown and is probably not true.

I don't think that models have been portrayed as equivalent to humans. Like most AI in it has been shown as vastly superior in some areas and profoundly ignorant in others. Media can overblow things and enthusiasts can talk about future advances as if they have already arrived, but I don't think these are typical portayals by the AI Field in general.

Re: Vision language models are blind

#107
post #106

Earlier quoted context omitted.

There are quite a few "ai apologists" in the comments but I think the title is fair when these models are marketed towards low vision people ("Be my eyes" https://www.youtube.com/watch?v=Zq710AKC1gg ) as the equivalent to human vision. These models are implied to be human level equivalents when they are not. This paper demonstrates that there are still some major gaps where simple problems confound the models in unex…

I disagree. I think the title, abstract, and conclusion not only misrepresents the state of the models but it misrepresents Thier own findings. They have identified a class of problems that the models perform poorly at and have given a good description of the failure. They portray this as a representative example of the behaviour in general. This has not been shown and is probably not true. I don't think that models…

Exactly... I've found GPT-4o to be good at OCR for instance... doesn't seem "blind" to me.

Re: Vision language models are blind

#108
post #106

Earlier quoted context omitted.

I disagree. I think the title, abstract, and conclusion not only misrepresents the state of the models but it misrepresents Thier own findings. They have identified a class of problems that the models perform poorly at and have given a good description of the failure. They portray this as a representative example of the behaviour in general. This has not been shown and is probably not true. I don't think that models…

Exactly... I've found GPT-4o to be good at OCR for instance... doesn't seem "blind" to me.

[deleted]

Re: Vision language models are blind

#109
post #65

I had a remarkable experience with GPT-4o yesterday. Our garage door started to fall down recently, so I inspected it and found that our landlord had installed the wire rope clips incorrectly, leading to the torsion cables losing tension. I didn't know what that piece of hardware was called, so I asked ChatGPT and it identified the part as I expected it to. As a test, I asked if there was anything notable about the p…

As a human, I was unable to see enough in that picture to infer which side was supposed to be under tension. I’m not trained, but I know what I expected to see from your description.

Like my sister post, I’m skeptical that the LLM didn’t just get lucky.

Re: Vision language models are blind

#110
post #69
post #65

I had a remarkable experience with GPT-4o yesterday. Our garage door started to fall down recently, so I inspected it and found that our landlord had installed the wire rope clips incorrectly, leading to the torsion cables losing tension. I didn't know what that piece of hardware was called, so I asked ChatGPT and it identified the part as I expected it to. As a test, I asked if there was anything notable about the p…

A human would need to trace the cable. An LLM may just be responding based on (1) the fact that you're asking about the clip in the first place, and that commonly happens when there's something wrong; and (2) that this is a very common failure mode. This is supported by it bringing up the "never saddle a dead horse" mnemonic, which suggests the issue is common. After you fix it, you should try asking the same questio…

[deleted]
Post reply on HN