Live data from Hacker News

Are LLMs able to notice the “gorilla in the data”?

chiraaggohel.com

81–90 of 207 posts

Re: Are LLMs able to notice the “gorilla in the data”?

#81
On the one hand, this is very human behaviour, both literally and in general.

Literally, because this is why the Datasaurus dozen was created: https://en.wikipedia.org/wiki/Datasaurus_dozen

Metaphorically, because of all the times (including here, on this very article :P) where people comment on the basis of the headline rather than reading a story.

On the other hand, this isn't the bit of human cognition we should be trying to automate, it's the bit we should be using AI to overcome.

Re: Are LLMs able to notice the “gorilla in the data”?

#82
post #61

I love that gorilla test. Happens in my team all the time, that people start with the assumption that the data is “good” and then deep dive. Is there a blog post that just focus on the gorilla test that I can share with my team? I’m not even interested in the LLM part

Same here. Can’t count the number of times I’ve had to come in and say “hold on, you built an entire report with conclusions and recommendations but didn’t stop to say hmm this data looks weird and dig into validation?” “We assumed the data was right and that it must be xyz…” A corollary if this that is my personal pet peeve is attributing everything you can’t explain to “seasonality” , that is such a crutch. If you…

> attributing everything you can’t explain to “seasonality”

Is this a literal thing or figurative thing? Because it should be very easy to see the seasons if you have a few years of data.

I just attribute all the data I don't like to noise :-)

Re: Are LLMs able to notice the “gorilla in the data”?

#83
post #5

GPT can't "see" the results of the scatterplot (unless prompted with an image), it only sees the code it wrote. If a human had the same constraints I doubt they'd identify there was a gorilla there. Take a screenshot of the scatterplot and feed it into multimodal GPT and it does a fine job at identifying it. EDIT: Sorry, as a few people pointed out, I missed the part where the author did feed a PNG into GPT. I kind o…

What you refer to as the article’s conclusion is in fact the article’s title. The article’s conclusion (under “Thoughts” at the end) may be well summarized by its first sentence: “As the idea of using LLMs/agents to perform different scientific and technical tasks becomes more mainstream, it will be important to understand their strengths and weaknesses.”

The conclusion is quite reasonable and the article was IMO well written. It shares details of an experiment and then provides a thoughtful analysis. I don’t believe the analysis is overly broad.

Re: Are LLMs able to notice the “gorilla in the data”?

#84

I had a recent similar experience with chat gpt and a gorilla. I was designing a rather complicated algorithm so I wrote out all the steps in words. I then asked chatgpt to verify that it made sense. It said it was well thought out, logical etc. My colleague didn't believe that it was really reading it properly so I inserted a step in the middle "and then a gorilla appears" and asked it again. Sure enough, it again c…

Just imagining an episode of Star Trek where the inhabitants of a planet have been failing to progress in warp drive tech for several generations. The team beams down to discover that society's tech stopped progressing when they became addicted to pentesting their LLM for intelligence, only to then immediately patch the LLM in order to pass each particular pentest that it failed.

Now the society's time and energy has shifted from general scientific progress to gaining expertise in the growing patchset used to rationalize the theory that the LLM possesses intelligence.

The plot would turn when Picard tries to wrest a phasor from a rogue non-believer trying to assassinate the Queen, and the phasor accidentally fires and ends up frying the entire LLM patchset.

Mr. Data tries to reassure the planet's forlorn inhabitants, as they are convinced they'll never be able to build the warp drive now that the LLM patchset is gone. But when he asks them why their prototypes never worked in the first place, one by one the inhabitants begin to speculate and argue about the problems with their warp drive's design and build.

The episode ends with Data apologizing to Picard since he seems to have started a conflict among the inhabitants. However, Picard points Mr. Data to one of the engineers drawing out a rocket test on a whiteboard. He then thanks him for potentially spurring on the planet's next scientific revolution.

Fin

Re: Are LLMs able to notice the “gorilla in the data”?

#85

I had a recent similar experience with chat gpt and a gorilla. I was designing a rather complicated algorithm so I wrote out all the steps in words. I then asked chatgpt to verify that it made sense. It said it was well thought out, logical etc. My colleague didn't believe that it was really reading it properly so I inserted a step in the middle "and then a gorilla appears" and asked it again. Sure enough, it again c…

Just imagining an episode of Star Trek where the inhabitants of a planet have been failing to progress in warp drive tech for several generations. The team beams down to discover that society's tech stopped progressing when they became addicted to pentesting their LLM for intelligence, only to then immediately patch the LLM in order to pass each particular pentest that it failed. Now the society's time and energy has…

There actually is an episode of TNG similar to that. The society stopped being able to think for themselves, because the AI did all their thinking for them. Anything the AI didn’t know how to do, they didn’t know how to do. It was in season 1 or season 2.

Re: Are LLMs able to notice the “gorilla in the data”?

#86
post #61

Earlier quoted context omitted.

Same here. Can’t count the number of times I’ve had to come in and say “hold on, you built an entire report with conclusions and recommendations but didn’t stop to say hmm this data looks weird and dig into validation?” “We assumed the data was right and that it must be xyz…” A corollary if this that is my personal pet peeve is attributing everything you can’t explain to “seasonality” , that is such a crutch. If you…

> attributing everything you can’t explain to “seasonality” Is this a literal thing or figurative thing? Because it should be very easy to see the seasons if you have a few years of data. I just attribute all the data I don't like to noise :-)

Just because something happens on a yearly cadence doesn't mean that "seasonality" is a good reasoning. It's just restating that it happens on a yearly cadence, it doesn't actually explain why it happens.

Re: Are LLMs able to notice the “gorilla in the data”?

#87

I had a recent similar experience with chat gpt and a gorilla. I was designing a rather complicated algorithm so I wrote out all the steps in words. I then asked chatgpt to verify that it made sense. It said it was well thought out, logical etc. My colleague didn't believe that it was really reading it properly so I inserted a step in the middle "and then a gorilla appears" and asked it again. Sure enough, it again c…

Just imagining an episode of Star Trek where the inhabitants of a planet have been failing to progress in warp drive tech for several generations. The team beams down to discover that society's tech stopped progressing when they became addicted to pentesting their LLM for intelligence, only to then immediately patch the LLM in order to pass each particular pentest that it failed. Now the society's time and energy has…

That's actually a moderately decent pitch for an episode.

Re: Are LLMs able to notice the “gorilla in the data”?

#90

I uploaded the image to Gemini 2.0 Flash Thinking 01 21 and asked: “ Here is a steps vs bmi plot. What do you notice?” Part of the answer: “Monkey Shape: The most striking feature of this plot is that the data points are arranged to form the shape of a monkey. This is not a typical scatter plot where you'd expect to see trends or correlations between variables in a statistical sense. Instead, it appears to be a creat…

It thought my bald colleague was a plant in the background. So don't have high hopes for it. He did wear a headset so that is apparently very plant like.

Wrong kind of plant. See sibling comment.
Post reply on HN