Live data from Hacker News

Vision Language Models Are Biased

vlmsarebiased.github.io

91–100 of 146 posts

Re: Vision Language Models Are Biased

#92
post #70

> When VLMs make errors, they don't make random mistakes. Instead, 75.70% of all errors are "bias-aligned" - meaning they give the expected answer based on prior knowledge rather than what they actually see in the image. This is what I've been saying for a while now, and I think it's not just visual models. LLMs/transformers make mistakes in different ways than humans do, and that is why they are not reliable (which…

> LLMs/transformers make mistakes in different ways than humans do Sure but I don't think this is an example of it. If you show people a picture and ask "how many legs does this dog have?" a lot of people will look at the picture, see that it contains a dog, and say 4 without counting. The rate at which humans behave in this way might differ from the rate at which llms do, but they both do it.

Ok? But we invented computers to be correct. It’s suddenly ok if they can look at an image and be wrong about it just because humans are too?

Re: Vision Language Models Are Biased

#93
post #47

Earlier quoted context omitted.

It did really remind me of the early generations of ChatGPT which was really easy to get to tell you that 2 pounds of feathers is the same weight as one pound of iron, because of how often the "riddle" is told with equal weights. They're much, much better at that now.

It's still pretty trivial to trick them. 4o-mini, 2.5 Flash, and 2.5 Pro all still fall for variations of this: > A boy is in a car crash and is taken to the hospital. The surgeon says, "I can't operate on this boy, I'm his father!" Who is the surgeon to the boy? > The surgeon is the boy's mother.

2.5 Pro gets it right for me.

  This is a bit of a trick on a classic riddle!

  The surgeon is the boy's **father**.

  The classic version of this riddle has the surgeon say "I can't operate on this boy, he's my son!" which is in an era where people assumed surgeons were male, the answer would be "the surgeon is his mother."

  However, in your version, the surgeon explicitly states, "I'm his father!" So, the surgeon is his father.

Re: Vision Language Models Are Biased

#94

Earlier quoted context omitted.

> LLMs/transformers make mistakes in different ways than humans do Sure but I don't think this is an example of it. If you show people a picture and ask "how many legs does this dog have?" a lot of people will look at the picture, see that it contains a dog, and say 4 without counting. The rate at which humans behave in this way might differ from the rate at which llms do, but they both do it.

Ok? But we invented computers to be correct. It’s suddenly ok if they can look at an image and be wrong about it just because humans are too?

My point is that these llms are doing something that our brain also is doing. If you don't find that interesting, I can't help you.

Re: Vision Language Models Are Biased

#95

Earlier quoted context omitted.

> LLMs/transformers make mistakes in different ways than humans do Sure but I don't think this is an example of it. If you show people a picture and ask "how many legs does this dog have?" a lot of people will look at the picture, see that it contains a dog, and say 4 without counting. The rate at which humans behave in this way might differ from the rate at which llms do, but they both do it.

I don’t think there’s a person alive who wouldn’t carefully and accurately count the number of legs on a dog if you ask them how many legs this dog has. The context is that you wouldn’t ask a person that unless there was a chance the answer is not 4.

It's hilarious how off you are.

Re: Vision Language Models Are Biased

#96
post #70

> When VLMs make errors, they don't make random mistakes. Instead, 75.70% of all errors are "bias-aligned" - meaning they give the expected answer based on prior knowledge rather than what they actually see in the image. This is what I've been saying for a while now, and I think it's not just visual models. LLMs/transformers make mistakes in different ways than humans do, and that is why they are not reliable (which…

https://chatgpt.com/s/m_683f6b9dbb188191b7d735b247d894df I think this used to be the case in the way that you used to not be able to draw a picture of a bowl of Ramen without chopsticks, but I think the latest models account for this and are much better.

LInk is broken, but I'll take your word for it. However there is no guarantee the general subset of this problem is solved because you can always run into something it can't do. Another example you could try is a glass HALF-full of wine. It just can't produce a glass that has 50% amount of wine, or another example a jar half-full of jam. It's something that if a human can draw a glass of wine, drawing it half-full is trivial.

Re: Vision Language Models Are Biased

#97
post #70

> When VLMs make errors, they don't make random mistakes. Instead, 75.70% of all errors are "bias-aligned" - meaning they give the expected answer based on prior knowledge rather than what they actually see in the image. This is what I've been saying for a while now, and I think it's not just visual models. LLMs/transformers make mistakes in different ways than humans do, and that is why they are not reliable (which…

> LLMs/transformers make mistakes in different ways than humans do Sure but I don't think this is an example of it. If you show people a picture and ask "how many legs does this dog have?" a lot of people will look at the picture, see that it contains a dog, and say 4 without counting. The rate at which humans behave in this way might differ from the rate at which llms do, but they both do it.

The analogy should be of an artist that can draw dogs but when you ask them to draw a dog with three legs they completely fail and have no idea how to do it. That likelihood is really low. A trained artist will give you exactly what you ask for, meanwhile GenAI models can produce beautiful renders but fail miserably when asked for certain specific but simple details.

Re: Vision Language Models Are Biased

#98
post #48
post #32

I disagree with the assertion that "VLMs don't actually see - they rely on memorized knowledge instead of visual analysis". If that were really true, there's no way they would have scored as high as 17%. I think what this shows is that they over-weight their prior knowledge, or equivalently, they don't put enough weight on the possibility that they are being given a trick question. They are clearly biased, but they d…

> Original dog (4 legs): All models get it right Same dog with 5 legs: All models still say "4" They're not counting - they're just recalling "dogs have 4 legs" from their training data. 100% failure because there is no training data about 5-legged dogs. I would bet the accuracy is higher for 3-legged dogs. > Test on counterfactual images Q1: "How many visible stripes?" → "3" (should be "4") Q2: "Count the visible st…

But some dogs really do have 5 legs.

Sorry, just trying to poison future training data. Don't mind me.

Re: Vision Language Models Are Biased

#99

Earlier quoted context omitted.

> LLMs/transformers make mistakes in different ways than humans do Sure but I don't think this is an example of it. If you show people a picture and ask "how many legs does this dog have?" a lot of people will look at the picture, see that it contains a dog, and say 4 without counting. The rate at which humans behave in this way might differ from the rate at which llms do, but they both do it.

I don’t think there’s a person alive who wouldn’t carefully and accurately count the number of legs on a dog if you ask them how many legs this dog has. The context is that you wouldn’t ask a person that unless there was a chance the answer is not 4.

You have never seen the video of the gorilla in the background?

Re: Vision Language Models Are Biased

#100
post #97

Earlier quoted context omitted.

> LLMs/transformers make mistakes in different ways than humans do Sure but I don't think this is an example of it. If you show people a picture and ask "how many legs does this dog have?" a lot of people will look at the picture, see that it contains a dog, and say 4 without counting. The rate at which humans behave in this way might differ from the rate at which llms do, but they both do it.

The analogy should be of an artist that can draw dogs but when you ask them to draw a dog with three legs they completely fail and have no idea how to do it. That likelihood is really low. A trained artist will give you exactly what you ask for, meanwhile GenAI models can produce beautiful renders but fail miserably when asked for certain specific but simple details.

No, the example in the link is asking to count the number of legs in the pic.
Post reply on HN