Live data from Hacker News

Vision Language Models Are Biased

vlmsarebiased.github.io

121–130 of 146 posts

Re: Vision Language Models Are Biased

#121
post #102
post #70

> When VLMs make errors, they don't make random mistakes. Instead, 75.70% of all errors are "bias-aligned" - meaning they give the expected answer based on prior knowledge rather than what they actually see in the image. This is what I've been saying for a while now, and I think it's not just visual models. LLMs/transformers make mistakes in different ways than humans do, and that is why they are not reliable (which…

> When VLMs make errors, they don't make random mistakes. Instead, 75.70% of all errors are "bias-aligned" - meaning they give the expected answer based on prior knowledge rather than what they actually see in the image. Yeah, that's exactly what our paper said 5 years ago! They didn't even cite us :( "Measuring Social Biases in Grounded Vision and Language Embeddings" https://arxiv.org/pdf/2002.08911

I think social biases (e.g. angry black women stereotype) in your paper is different from cognitive biases about facts (e.g. number of legs, whether lines are parallel) that OP is about.

Social biases are subjective. Facts are not.

Re: Vision Language Models Are Biased

#122
post #99

Earlier quoted context omitted.

I don’t think there’s a person alive who wouldn’t carefully and accurately count the number of legs on a dog if you ask them how many legs this dog has. The context is that you wouldn’t ask a person that unless there was a chance the answer is not 4.

You have never seen the video of the gorilla in the background?

That's a specific example that when you draw a human's attention to something (eg: count the number of ball passes in this video), they hyper-fixate on that, to the exclusion of other things, so it seems like it makes the opposite point that I think you're trying to?

Re: Vision Language Models Are Biased

#123
post #102
post #70

> When VLMs make errors, they don't make random mistakes. Instead, 75.70% of all errors are "bias-aligned" - meaning they give the expected answer based on prior knowledge rather than what they actually see in the image. This is what I've been saying for a while now, and I think it's not just visual models. LLMs/transformers make mistakes in different ways than humans do, and that is why they are not reliable (which…

> When VLMs make errors, they don't make random mistakes. Instead, 75.70% of all errors are "bias-aligned" - meaning they give the expected answer based on prior knowledge rather than what they actually see in the image. Yeah, that's exactly what our paper said 5 years ago! They didn't even cite us :( "Measuring Social Biases in Grounded Vision and Language Embeddings" https://arxiv.org/pdf/2002.08911

What do you genuinely think they built upon from your paper?

If anything, the presentation of their results in such an accessible format next to the paper should be commended.

Re: Vision Language Models Are Biased

#124
post #102
post #70

> When VLMs make errors, they don't make random mistakes. Instead, 75.70% of all errors are "bias-aligned" - meaning they give the expected answer based on prior knowledge rather than what they actually see in the image. This is what I've been saying for a while now, and I think it's not just visual models. LLMs/transformers make mistakes in different ways than humans do, and that is why they are not reliable (which…

> When VLMs make errors, they don't make random mistakes. Instead, 75.70% of all errors are "bias-aligned" - meaning they give the expected answer based on prior knowledge rather than what they actually see in the image. Yeah, that's exactly what our paper said 5 years ago! They didn't even cite us :( "Measuring Social Biases in Grounded Vision and Language Embeddings" https://arxiv.org/pdf/2002.08911

Well you send a vaguely worded email like "I think you may find our work relevant" and everyone knows what that means and adds the citation

Re: Vision Language Models Are Biased

#125
post #102

Earlier quoted context omitted.

> When VLMs make errors, they don't make random mistakes. Instead, 75.70% of all errors are "bias-aligned" - meaning they give the expected answer based on prior knowledge rather than what they actually see in the image. Yeah, that's exactly what our paper said 5 years ago! They didn't even cite us :( "Measuring Social Biases in Grounded Vision and Language Embeddings" https://arxiv.org/pdf/2002.08911

I think social biases (e.g. angry black women stereotype) in your paper is different from cognitive biases about facts (e.g. number of legs, whether lines are parallel) that OP is about. Social biases are subjective. Facts are not.

As far as the model's concerned, there's not much difference. Social biases will tend to show up objectively in the training data because the training data is influenced by those biases (the same thing happens with humans, which how these biases can proliferate and persist).

Re: Vision Language Models Are Biased

#126

Earlier quoted context omitted.

I think social biases (e.g. angry black women stereotype) in your paper is different from cognitive biases about facts (e.g. number of legs, whether lines are parallel) that OP is about. Social biases are subjective. Facts are not.

As far as the model's concerned, there's not much difference. Social biases will tend to show up objectively in the training data because the training data is influenced by those biases (the same thing happens with humans, which how these biases can proliferate and persist).

I see a clear difference. One is objective (only one correct answer), one is subjective (multiple plausible answers)

Re: Vision Language Models Are Biased

#127

Unless the training set was explicitly biased in a specific way, this is basically saying that "the world is biased"

Models can be biased, but it doesn't seem like it should be a reason to get the answer wrong, right? Humans have biases too, but we don't get those simple questions wrong

Re: Vision Language Models Are Biased

#128

The basic results are interesting, but what really surprised me is that asking them to double-check didn't work. Falling for an "optical illusion" is one thing, but being unable to see the truth once you know the illusion there is much worse.

Because it isn’t thinking. Asking it to “double check” is like pressing the equals button on a calculator a second time. It just runs the same calculation again.

Re: Vision Language Models Are Biased

#129
post #14

FWIW I tried the first couple of examples in ChatGPT 4o and couldn't replicate this. For example: "The animal in the image is a chicken, and it appears to have four legs. However, chickens normally have only two legs. The presence of four legs suggests that the image may have been digitally altered or artificially generated." I don't have a good explanation for why I got different results.

I suspect that responses are altered/corrected based on what people query from popular online models. I have had several occasions that I ask some "How do I ... in X software?" question some day and model keeps hallucinating nonexistant config options regardless how many times I keep saying "This option doesn't exist in software X". But if I asked the same question some days later, the answer was completely different and made even some sense.

Re: Vision Language Models Are Biased

#130
post #61
post #51

Earlier quoted context omitted.

> If you "saw" [a three-legged chicken] in real life, you'd probably rub your eyes and discount it too. Huh? I'd assume it's a mutant, not store a memory of having seen a perfectly normal chicken You've never seen someone who's missing a finger or has only a half-grown arm or something? Surely you didn't assume your eyes were tricking you?! Or... if you did, I guess you can't answer this question. I'm actually rackin…

You've seen people with missing limbs without being surprised, because you know how they can become lost, but you rarely see one with additional limbs. Their likelihoods and our consequent priors are drastically different. Also, your reaction will depend on how strong the evidence is. Did you 'see' the three-legged chicken pass by some bush in the distance, or was it right in front of you?

But to be clear, in this case the LLM has a full, direct, unobscured view of the chicken. A human, in that specific case -- i.e. looking at the same photo* -- would not have trouble discerning and reporting the third leg. Perhaps if they were forced to scan the photo quickly and make a report, or were otherwise not really 'paying attention'/'taking it seriously', but the mere fact that LLMs fall into that regime far more than an 'serious employee' already shows that they fail in different ways than humans do.
Post reply on HN