Live data from Hacker News

Are LLMs able to notice the “gorilla in the data”?

chiraaggohel.com

181–190 of 207 posts

Re: Are LLMs able to notice the “gorilla in the data”?

#181
Am I the only one who's more shocked by the LLMs affirming "The distributions appear roughly normal for both genders, as shown in the visualization", "Both distributions appear approximately normal, though with some right skew" and such than by any gorilla issue?

From short thinking or from looking at the graphs I would believe "roughly normal" sounds like wishful thinking to stay in the reassuring bounds of normal distributions. And I believe things would get dangerous once you would start using these assumptions for tests and affirmations.

My short thinking: distributions don't look close to normal on the graphs. Values are probably bounded on one side and almost unbounded on the other (can't go below 0 steps, can go into very high number of steps on 1 day). There are days / people with close to 0 steps and others that might distribute in a sort of normal around a value maybe. Weight and height might be normally distributed in a population but they're correlated and BMI is one divided by the square of the other. I can't compute the resulting distribution but I would doubt that would make for a distribution close to normal.

Ok the LLMs were told to assume both traits were distributed normally, but affirming they look mostly normal is scary to me.

Am I too picky and in real analyses assuming such distributions are "mostly normal" is fine for all practical purposes?

Re: Are LLMs able to notice the “gorilla in the data”?

#182

Am I the only one who's more shocked by the LLMs affirming "The distributions appear roughly normal for both genders, as shown in the visualization", "Both distributions appear approximately normal, though with some right skew" and such than by any gorilla issue? From short thinking or from looking at the graphs I would believe "roughly normal" sounds like wishful thinking to stay in the reassuring bounds of normal d…

Honestly, this was the meta-gorilla in the data for me! I was so busy focusing on the LLM’s EDA that I didn’t really interrogate some of the other data analysis practices.

In general, I’ve steered clear of current LLMs for data analysis/description because they seem so highly influenced by choice of prompt and wording. They tend to simply affirm any language I use to describe the data initially.

To be fair, I’ve attended conferences and lab meetings where humans will refer to a any vaguely concave curved distribution as “mostly normal” :P

Re: Are LLMs able to notice the “gorilla in the data”?

#183

Earlier quoted context omitted.

> That there was a revolt against thinking machines... Yes... > ...because humans had become too dependent on them for thinking. ... but no. The causes of the Butlerian Jihad are forgotten (or, at least, never mentioned) in any of Frank Herbert's novels; all that's remembered is the outcome.

>> ...because humans had become too dependent on them for thinking. > ... but no. The causes of the Butlerian Jihad are forgotten (or, at least, never mentioned) in any of Frank Herbert's novels; all that's remembered is the outcome. Per Wikipedia or Goodreads, God Emperor of Dune has "The target of the Jihad was a machine-attitude as much as the machines...Humans had set those machines to usurp our sense of beauty,…

It's still a little ambiguous - and perhaps deliberately so - whether Leto is describing what inspired the Jihad, or what it became. The series makes it quite clear that the two are often not the same. As Leto continues later in that chapter:

"Throughout our history, the most potent use of words has been to round out some transcendental event, giving that event a place in the accepted chronicles, explaining the event in such a way that ever afterward we can use those words and say: 'This is what it meant.' That's how events get lost in history."

Re: Are LLMs able to notice the “gorilla in the data”?

#184
post #45
post #33

Earlier quoted context omitted.

The title of the article is "Your AI Can't See Gorillas". That seems demonstrably false. The article says: > Furthermore, their data analysis capabilities seem to focus much more on quantitative metrics and summary statistics, and less on the visual structure of the data Again, this seems false - or, at best, misleading. I had no problem getting AI to focus on visual structure of the data without any tricks. A more f…

you knew that there was a visual gag in there before asking it to. if you didnt know it was there, and took a look at only the text output, the llm would not have found it to tell you its there

Yeah the people "it gives you the answer when you give it the answer" have kind of ruined my morning. Oh well.

Re: Are LLMs able to notice the “gorilla in the data”?

#185

Earlier quoted context omitted.

The Asimov story it reminded me of was The Profession, though that one is not really about AI - but it is about original ideas and the kinds of people that have them. I find the LLM dismissals somewhat tedious for most of the people making them half of humanity wouldn't meet their standards.

> I find the LLM dismissals somewhat tedious for most of the people making them half of humanity wouldn't meet their standards. Aren't people funny like that? One person values an encyclopedic chatbot for company, the next prefers a human. Thank god we can all get along.

That isn't what I said, but likely it won't matter. People will be denying it up until the end. I don't prefer LLMs to humans, but I don't pretend biological minds contain some magical essence that separates us from silicon. The denials of what might be happening are pretty weak - at best they're way over confident and smug.

Re: Are LLMs able to notice the “gorilla in the data”?

#186

Earlier quoted context omitted.

I think it depends if one is using “AI” as a tool or as a replacement for an intelligent expert? The former, sure, it’s maybe not expected, because the prompter is already an intelligent expert. If the latter, then yes, I think, because if you gave the task to an expect and they did not notice this, I would consider them not good at their job. See also Anscombe's quartet[1] and the Datasaurus dozen[2] (mentioned in a…

This is true, but I would replace 'intelligent expert' with 'intelligent human expert'. Graphing data to analyze it - and then seeing shapes and creatures in said graph - is a distinctly human practice, and not an inherently necessary part of most data analysis (the obvious exception being when said data draws a picture). I think it's because the interface uses human language that people expect AI to make the same as…

  > not an inherently necessary part of most data analysis
You do realize that the LLMs did not find the data suspicious, right? I think your answer is appropriate if they answered (without follow-up prompting which is leaking information to the LLM!) that the data was suspicious. But in fact, all models are saying that the data is normally distributed. Sure, the author said this, but they confirmed it. If you run normaltest on any BMI or steps, you'll find that they are very NOT normal. In fact, you can also see this from the histograms.

So honestly, this isn't even about the Gorilla. You're hyper focused there because you're looking for a way to make the LLM right while not looking for why the LLM got it wrong (it did, there's no denying it, so we should understand why it is wrong, right?). The problem isn't so much about expecting it to be human, the problem is if it can do data analysis. The problem here is that the LLM will not correct you, it will not "trust but verify" you. It is a "yes man" and is trained to generate outputs that optimize human preference. That last part alone should make you extremely suspicious, as it means when it is wrong, it is more likely to be in exactly the way you won't notice.

Re: Are LLMs able to notice the “gorilla in the data”?

#187

Earlier quoted context omitted.

Just imagining an episode of Star Trek where the inhabitants of a planet have been failing to progress in warp drive tech for several generations. The team beams down to discover that society's tech stopped progressing when they became addicted to pentesting their LLM for intelligence, only to then immediately patch the LLM in order to pass each particular pentest that it failed. Now the society's time and energy has…

There actually is an episode of TNG similar to that. The society stopped being able to think for themselves, because the AI did all their thinking for them. Anything the AI didn’t know how to do, they didn’t know how to do. It was in season 1 or season 2.

Wondering if you're thinking of some TOS episodes such as "Taste of Armageddon" or "The Ultimate Computer" (M5)?

Re: Are LLMs able to notice the “gorilla in the data”?

#188
post #11
post #7

Can it draw the unicorn yet? https://gpt-unicorn.adamkdean.co.uk/

Claude 3.5 Sonnet is much better at it: https://claude.site/artifacts/ad1b544f-4d1b-4fc2-9862-d6438e... But I guess GPT-4o results are more funny to look at.

Two pink squiggles is better than the hundred listed in the OP link?

Re: Are LLMs able to notice the “gorilla in the data”?

#189

Earlier quoted context omitted.

> I find the LLM dismissals somewhat tedious for most of the people making them half of humanity wouldn't meet their standards. Aren't people funny like that? One person values an encyclopedic chatbot for company, the next prefers a human. Thank god we can all get along.

That isn't what I said, but likely it won't matter. People will be denying it up until the end. I don't prefer LLMs to humans, but I don't pretend biological minds contain some magical essence that separates us from silicon. The denials of what might be happening are pretty weak - at best they're way over confident and smug.

[deleted]

Re: Are LLMs able to notice the “gorilla in the data”?

#190

Earlier quoted context omitted.

Without an image? No, not at all. It's supposed to make its own image. And it did make its own image. But it didn't properly analyze the image it made.

That's a feature that would need to be implemented. There's no reason to think it could look at the image of the plot it generated automatically, but feeding it the image it generated back to it is no different to if it did view it automatically

The point of telling it to explore the data is so I don't have to think of every angle myself. Humans can get an understanding from visuals that LLMs can't match, apparently, even without gimmicks.
Post reply on HN