Live data from Hacker News

Are LLMs able to notice the “gorilla in the data”?

chiraaggohel.com

31–40 of 207 posts

Re: Are LLMs able to notice the “gorilla in the data”?

#31
post #19
post #15

Earlier quoted context omitted.

Before seeing Claude’s response, did you see where the author said > I asked the model to closely look at the plot, and also uploaded a png of the plot it had generated.

I sent the plot to ChatGPT 4o. Here is the conversation: what do you see ChatGPT said: This is a scatter plot with the variables "steps" on the x-axis and "bmi" on the y-axis. The data points are colored by "gender" (red for female and blue for male). Interestingly, the arrangement of the points appears to form a drawing resembling a cartoonish figure or character, likely added for artistic or humorous effect. If you…

Sure whatever.

OC seemed to think that Claude did that with just the data and not the image of the scatterplot it’s.

Re: Are LLMs able to notice the “gorilla in the data”?

#32
post #10

Earlier quoted context omitted.

That has nothing to do with this article. Wrong gorilla.

[flagged]

it doesn't, it just means that "AI gorilla controversy" was a big deal not too long ago, and they happened to remember it.

Re: Are LLMs able to notice the “gorilla in the data”?

#33
post #21

Earlier quoted context omitted.

Hm, interesting. The way I tried it was by pasting an image into Claude directly as the start of the conversation, plus a simple prompt ("What do you see here?"). It got the specific image wrong (it thought it was baby yoda, lol), but it did understand that it was an image. I wonder if the author got different results because they had been talking a lot about a data set before showing the image, which possibly predis…

Please read TFA. The conclusion of the article isn't nearly so simplistic, they're just suggesting that you have to be aware of the natural strengths and weaknesses of LLMs, even multi modal ones particularly around visual pattern recognition vs quantitative pattern recognition. And yes, the idea that the initial context can sometimes predispose the LLM to consider things in a more narrow manner than a user might oth…

The title of the article is "Your AI Can't See Gorillas". That seems demonstrably false.

The article says:

> Furthermore, their data analysis capabilities seem to focus much more on quantitative metrics and summary statistics, and less on the visual structure of the data

Again, this seems false - or, at best, misleading. I had no problem getting AI to focus on visual structure of the data without any tricks. A more fair statement would be "If you ask an AI a bunch of questions about summary statistics and then show it a scatterplot with an image, then it might continue to focus on summary statistics". But that's not what the concluding paragraph states, and it's not what the title states, either.

Re: Are LLMs able to notice the “gorilla in the data”?

#36
post #31
post #19

Earlier quoted context omitted.

I sent the plot to ChatGPT 4o. Here is the conversation: what do you see ChatGPT said: This is a scatter plot with the variables "steps" on the x-axis and "bmi" on the y-axis. The data points are colored by "gender" (red for female and blue for male). Interestingly, the arrangement of the points appears to form a drawing resembling a cartoonish figure or character, likely added for artistic or humorous effect. If you…

Sure whatever. OC seemed to think that Claude did that with just the data and not the image of the scatterplot it’s.

LLM responses are random. One's failure is other's success. When evaluating we all should do rerurns and see how many times it fails or succeeds.

Without number of rerurns, the result is as good as random.

Re: Are LLMs able to notice the “gorilla in the data”?

#38
If you give a blind researcher this task, they might have trouble seeing the gorillas as well.

Also the prompt matters. To a human, literally everything they see and experience is "the prompt", so to speak. A constant barrage of inputs.

To the AI, it's just the prompt and the text it generates.

Re: Are LLMs able to notice the “gorilla in the data”?

#40
I love that gorilla test. Happens in my team all the time, that people start with the assumption that the data is “good” and then deep dive.

Is there a blog post that just focus on the gorilla test that I can share with my team? I’m not even interested in the LLM part

Post reply on HN