Live data from Hacker News

Are LLMs able to notice the “gorilla in the data”?

chiraaggohel.com

111–120 of 207 posts

Re: Are LLMs able to notice the “gorilla in the data”?

#111
This doesn't seem to make sense. Can a human spot a gorilla in a sequence of numbers? Try it. Later on, he gives it a picture and it correctly spots the mistake.

>but does not specifically understand the pattern as a gorilla

Maybe it does, how could you tell? Do you really expect an assistant to say "Holy shit, there's a gorilla in your plot!"? The only thing relevant to the request is that the data seems fishy, and it outputs exactly this. Maybe something trained for creative writing, agency, character, and witty remarks (like Claude 3 Opus) would be inclined to do that, and that would be amusing, but that's pretty optional for the presented task.

Re: Are LLMs able to notice the “gorilla in the data”?

#112

These posts about X task LLMs fails at when you give it Y prompt are getting more and more silly. If you ask an AI to analyze some data, should the default behavior be to use that data to make various types of graphs, export said graphs, feed them back in to itself, then analyze the shapes of those graphs to see if they resemble an animal? Personally I would be very annoyed if I actually wanted a statistical analysis…

I think it depends if one is using “AI” as a tool or as a replacement for an intelligent expert? The former, sure, it’s maybe not expected, because the prompter is already an intelligent expert. If the latter, then yes, I think, because if you gave the task to an expect and they did not notice this, I would consider them not good at their job. See also Anscombe's quartet[1] and the Datasaurus dozen[2] (mentioned in another comment as well).

[1]: https://en.wikipedia.org/wiki/Anscombe's_quartet [2]: https://en.wikipedia.org/wiki/Datasaurus_dozen

Re: Are LLMs able to notice the “gorilla in the data”?

#113

I uploaded the image to Gemini 2.0 Flash Thinking 01 21 and asked: “ Here is a steps vs bmi plot. What do you notice?” Part of the answer: “Monkey Shape: The most striking feature of this plot is that the data points are arranged to form the shape of a monkey. This is not a typical scatter plot where you'd expect to see trends or correlations between variables in a statistical sense. Instead, it appears to be a creat…

That doesn't seem like an appropriate comparison to the task the blogger did. The blogger gave their AI thing the raw data - and a different prompt from the one you gave. If you gave it a raster image, that's "cheating" - these models were trained to recognize things in images.

In the article they also give the png

Re: Are LLMs able to notice the “gorilla in the data”?

#115

maybe i dont get it, but can we conclusively say that the gorilla wasn’t “seen” vs. deemed to be irrelevant to the questions being asked? “look at the scatter plot again” is anthropomorphizing the llm and expecting it to infer a fairly odd intent. would queries like, “does the scatter plot visualization look like any real world objects?” may have produced a result the author was fishing for. if it were the opposite s…

If you were trying to answer real questions, you’d want to know if there were clear signs of the data being fake, flawed, or just different-looking than expected, potentially leading to new hypotheses.

The gorilla is just an extreme example of that.

Albeit perhaps an unfair example when applied to AI.

In the original experiment with humans, the assumption seemed to be that the gorilla is fundamentally easy to see. Therefore if you look at the graph to try to find patterns in it, you ought to notice the gorilla. If you don’t notice it, you might also fail to notice other obvious patterns that would be more likely to occur in real data.

Even for humans, that assumption might be incorrect. To some extent, failing to notice the gorilla might just be demonstrating a quirk in our brains’ visual processing. If we expect data, we see data, no matter how obvious the gorilla might be. Failing to notice the gorilla doesn’t necessarily mean that we’d also fail to notice the sorts of patterns or flaws that appear in real data. But on the other hand, people do often fail to notice ‘obvious’ patterns in real data. To distinguish the two effects, you’d want a larger experiment with more types of ‘obvious’ flaws than just gorillas.

For AI, those concerns are the same but magnified. On one hand, vision models are so alien that it’s entirely plausible they can notice patterns reliably despite not seeing the gorilla. On the other hand, vision models are so unreliable that it’s also plausible they can’t notice patterns in graphs well at all.

In any case, for both humans and AI, it’s interesting what these examples reveal about their visual processing, which is in both cases something of a black box. That makes the gorilla experiment worth talking about regardless of what lessons it does or doesn’t hold for real data analysis.

Re: Are LLMs able to notice the “gorilla in the data”?

#116
post #31

Earlier quoted context omitted.

Sure whatever. OC seemed to think that Claude did that with just the data and not the image of the scatterplot it’s.

LLM responses are random. One's failure is other's success. When evaluating we all should do rerurns and see how many times it fails or succeeds. Without number of rerurns, the result is as good as random.

Okay?

OC was saying that the article said that Claude recognized the “artistic” lines of the image from just the scatter plot data.

That isn’t what happened.

The author added a png of the plot to the conversation.

Idk why I need to explain that twice.

Re: Are LLMs able to notice the “gorilla in the data”?

#117
post #58
post #7

Can it draw the unicorn yet? https://gpt-unicorn.adamkdean.co.uk/

I wondered if o1 would do better- seems reasonable that step-by-step trying to produce legs/torso/head/horn would do better than very weird legless things 4o is making. Looks like someone has done it: https://openaiwatch.com/?model=o1-preview They do seem to generally have legs and head, which is an improvement over 4o. Still pretty unimpressive.

Why not o3-mini?

Re: Are LLMs able to notice the “gorilla in the data”?

#118

Earlier quoted context omitted.

That doesn't seem like an appropriate comparison to the task the blogger did. The blogger gave their AI thing the raw data - and a different prompt from the one you gave. If you gave it a raster image, that's "cheating" - these models were trained to recognize things in images.

In the article they also give the png

> When a png is directly uploaded, the model is better able to notice that some strange pattern is present in the data. However, it still does not recognize the pattern as a gorilla.

I wonder if the conversation context unfairly weighed the new impression towards the previous interpretation.

Re: Are LLMs able to notice the “gorilla in the data”?

#120

Earlier quoted context omitted.

“The ball bearings are to be made of wood, because no one is going to read this work this far anyway.”

And a bowl of M&Ms, with all the brown ones taken out - to make sure they did read this far.

I still think that is such a simple, easy litmus test... genius :-)
Post reply on HN