Live data from Hacker News

Vision Language Models Are Biased

vlmsarebiased.github.io

101–110 of 146 posts

Re: Vision Language Models Are Biased

#101

Earlier quoted context omitted.

Ok? But we invented computers to be correct. It’s suddenly ok if they can look at an image and be wrong about it just because humans are too?

My point is that these llms are doing something that our brain also is doing. If you don't find that interesting, I can't help you.

Well, they’re getting the same result. I don’t particularly see why that’s useful.

Re: Vision Language Models Are Biased

#102
post #70

> When VLMs make errors, they don't make random mistakes. Instead, 75.70% of all errors are "bias-aligned" - meaning they give the expected answer based on prior knowledge rather than what they actually see in the image. This is what I've been saying for a while now, and I think it's not just visual models. LLMs/transformers make mistakes in different ways than humans do, and that is why they are not reliable (which…

> When VLMs make errors, they don't make random mistakes. Instead, 75.70% of all errors are "bias-aligned" - meaning they give the expected answer based on prior knowledge rather than what they actually see in the image.

Yeah, that's exactly what our paper said 5 years ago!

They didn't even cite us :(

"Measuring Social Biases in Grounded Vision and Language Embeddings" https://arxiv.org/pdf/2002.08911

Re: Vision Language Models Are Biased

#103

Earlier quoted context omitted.

My point is that these llms are doing something that our brain also is doing. If you don't find that interesting, I can't help you.

Well, they’re getting the same result. I don’t particularly see why that’s useful.

All automation has ever been is an object doing something that a human can do, without needing the human.

Re: Vision Language Models Are Biased

#104

Models are Bias A model is bias, implemented as a collection of statistics that weigh relationships between given tokens. It doesn't deduce or follow logic. It doesn't make or respect categories. It just shows you what in its data set is most familiar to what is in your prompt; where familiarity is defined implicitly by the makeup of the original training corpus, and explicitly by the training weights. We need to sto…

How are you defining “bias”?

The definition I’ve found useful (outside of the “the constant term contribution”) is “a tendency to be wrong in an identifiable direction”.

But that doesn’t seem to be the definition you are using. So, what do you mean?

Re: Vision Language Models Are Biased

#105
post #97

Earlier quoted context omitted.

The analogy should be of an artist that can draw dogs but when you ask them to draw a dog with three legs they completely fail and have no idea how to do it. That likelihood is really low. A trained artist will give you exactly what you ask for, meanwhile GenAI models can produce beautiful renders but fail miserably when asked for certain specific but simple details.

No, the example in the link is asking to count the number of legs in the pic.

Ok, sure, but I'm trying to point out the gap in expectation, i.e. it's an expert artist but it cannot fulfill certain specific but simple requests.

Re: Vision Language Models Are Biased

#106

Earlier quoted context omitted.

Well, they’re getting the same result. I don’t particularly see why that’s useful.

All automation has ever been is an object doing something that a human can do, without needing the human.

The result is still wrong, though! It needs to be right to be useful!

Re: Vision Language Models Are Biased

#107
post #96

Earlier quoted context omitted.

https://chatgpt.com/s/m_683f6b9dbb188191b7d735b247d894df I think this used to be the case in the way that you used to not be able to draw a picture of a bowl of Ramen without chopsticks, but I think the latest models account for this and are much better.

LInk is broken, but I'll take your word for it. However there is no guarantee the general subset of this problem is solved because you can always run into something it can't do. Another example you could try is a glass HALF-full of wine. It just can't produce a glass that has 50% amount of wine, or another example a jar half-full of jam. It's something that if a human can draw a glass of wine, drawing it half-full is…

chatgpt can easily do that? What was the last time you tried?

Re: Vision Language Models Are Biased

#108

Earlier quoted context omitted.

Annoying. The actual braille on the sign was "⠁⠒⠑⠎⠎⠊⠼" which I gather means "accessible" in abbreviated braille. None of my attempts got it to even transcribe it to Unicode characters properly. I got "elevator", "friend", etc. Just wildly making stuff up and completely useless, even when it wasn't distracted by the No Smoking sign (in the second case I cropped out the rest of the sign). And in all cases, supremely co…

> This seems like something a VLM should handle very easily Not if its training data doesn't include braille as first class but has lots of braille signage with bad description (e.g., because people assumed the accompanying English matches the braille.) This could very well be the kind of mundane AI bias problem that the x-risk and tell-me-how-to-make-WMD concerns have shifted concerns about problems in AI away from.

I'd wager that correctly labeled braille far exceeds dumb braille, and when presented with just the braille it flat out hallucinated braille characters that weren't there. It didn't seem to actually be parsing the dots at all. My theory is that it has hardly seen any braille, despite it insisting that it knows how to read it.

Re: Vision Language Models Are Biased

#110
post #61
post #51

Earlier quoted context omitted.

> If you "saw" [a three-legged chicken] in real life, you'd probably rub your eyes and discount it too. Huh? I'd assume it's a mutant, not store a memory of having seen a perfectly normal chicken You've never seen someone who's missing a finger or has only a half-grown arm or something? Surely you didn't assume your eyes were tricking you?! Or... if you did, I guess you can't answer this question. I'm actually rackin…

You've seen people with missing limbs without being surprised, because you know how they can become lost, but you rarely see one with additional limbs. Their likelihoods and our consequent priors are drastically different. Also, your reaction will depend on how strong the evidence is. Did you 'see' the three-legged chicken pass by some bush in the distance, or was it right in front of you?

There's a first time you see everything you don't know how to explain.
Post reply on HN