Live data from Hacker News

What Emily Bender meant by "stochastic parrots"

spectrum.ieee.org

241–250 of 280 posts

Re: What Emily Bender meant by "stochastic parrots"

#241
post #45

I'm sorry but I do tend to feel like this muddies up the discussion on "what this technology really is". I think "artificial" is actually a pretty good term to describe the output of the models. That output does appear to resemble at least some definition of the word "intelligence" - there is some ability there to do cognition over information that's been provided to them in-context. What is it to understand, then? I…

> a better term than artificial intelligence

Agreed, it's a problematic term that conflates theoretical research with better search results.

Deep learning and machine learning (ML) are both unromanticised (un-hyped) terms.

In some cases, it's just tooling with a better interface. As we have done with other complex computing systems (e.g. Deep Blue, Watson), we might just end up naming it a computing system for querying, e.g.

https://en.wikipedia.org/wiki/LCARS

We would like a neutral term for such a system, and in this sense it's better to call it "A.I." than to make a verb from google or bing or other corporate name.

But in refined cases, such tooling may offer a credible (as in "believable") human experience, like Weizenbaum's Eliza (https://en.wikipedia.org/wiki/ELIZA). Some users of Eliza fully believed that Eliza listened and understood at a profound human level. Eliza was a simple computer programme.

Computing systems are not humans. They have no accountability in real life. People may be so comfortable with the user-experience that they cannot distinguish it from interacting with another human. That doesn't make the computer system alive and accountable. And if it's run by a corporation, it will almost certainly make big promises while energetically seeking to avoid accountability. ^_^

Re: What Emily Bender meant by "stochastic parrots"

#242

Earlier quoted context omitted.

Oh, now I see where we have an actual difference of opinion. I don't think you can deny that even Finnegan's wake proceeds one token at a time; your interpretation of it may require more context or out-of-order interpretation, but that's just as true when observing text in German or Japanese, which have word ordering constraints that are alien to English speakers. How it was written is irrelevant; all we can observe…

I am no linguist, but I believe this is referred to as surface structure and deep structure. What you describe is a line of text that is grammatically somewhat adequate to pass as readable and you treat all text the same. when we read the text, we decipher meaning out of according to word references and syntax - just like with Python or C++. However, if this were solely the case, we could not read Finnegan’s Wake. Pr…

What I'm saying is that we cannot observe deep structure. All we can observe is the surface structure. If we have access to the author or speaker we can ask clarifying questions, etc., but even that just produces more examples of surface structure.

But even here, in the Chomsky sense, LLMs clearly exhibit deep structure because they can write at length in an internally consistent manner. Importantly, early generations, GPT-2 and even GPT-3, did not definitively have this property; roughly, an object that was green at the beginning of a paragraph might not still be green at the end of the paragraph. This was strong evidence for lack of a world model.

Current LLMs do not show this behavior. We cannot prove that LLMs have a world model, in fact, their architecture seems to rule it out, but looking at it from a linguistic standpoint, they produce language in a manner as if to reflect a world view. That is, we cannot easily falsify the statement "LLMs somehow represent a world model"; and current examples of "disproving" their world view are so convoluted that even humans do not appear to (observationally) have a world view either.

I'm not making a claim as to LLMs having genuine deep structure or consciousness or anything like that. I'm claiming that we can't rule out current or future capabilities or make structural assumptions. Yes, they generalize from their training data, but unless you can make very specific claims about the kinds of things that they _cannot_ do, I can't take this statement as particularly compelling.

Re: What Emily Bender meant by "stochastic parrots"

#243
post #197

Earlier quoted context omitted.

No, they were correct. In fact an LLM stitches together stuff it observed in its training data. That scales up way better than a lot of us expected, but it's still correct. If you train it on lots of working code, then it's useful for coding. If you trained it primarily on non-working code it would produce nonsense.

It's not correct. Please read some more research papers, this isn't what people working in AI believe at all. You can prove with experiments that different human languages get translated to the same abstract conceptual space in the middle layers, for example. It's why interpretability is so difficult. The claim is odd in another way: you can train a person on non-working code and they'll produce nonsense. That doesn'…

[deleted]

Re: What Emily Bender meant by "stochastic parrots"

#244
post #197

Earlier quoted context omitted.

No, they were correct. In fact an LLM stitches together stuff it observed in its training data. That scales up way better than a lot of us expected, but it's still correct. If you train it on lots of working code, then it's useful for coding. If you trained it primarily on non-working code it would produce nonsense.

It's not correct. Please read some more research papers, this isn't what people working in AI believe at all. You can prove with experiments that different human languages get translated to the same abstract conceptual space in the middle layers, for example. It's why interpretability is so difficult. The claim is odd in another way: you can train a person on non-working code and they'll produce nonsense. That doesn'…

[deleted]

Re: What Emily Bender meant by "stochastic parrots"

#245

Earlier quoted context omitted.

you are right, i was more curt than i should have been. apologies. but you helped prove my point: >>"This sentence has five words" is going to appear far more often than "This sentence has four words". it's not about this at all. your point is about data quality. you need to take a step back. the point is that if you trained a language model just on this data set which has sentences akin to "this sentence has two wor…

If someone substituted all of your sensory inputs for something else for your entire life, how would you notice? If you wore contacts from birth that made the sky red and earbuds that censored when people said it was blue, on what basis would you realize that was wrong? I don't see what this says about the architecture of your brain, and I don't think it's the point being made in the paper. That the training data mus…

>>If someone substituted all of your sensory inputs for something else for your entire life, how would you notice?

I don't know how or if I will notice. That's the biology, chemistry and physics of the brain that I don't know. I hope someone is looking into it. But this does not mean in any way that LLMs are similar to our brains!!!!!! WE DON'T KNOW HOW OUR BRAINS WORK. So going back to language modeling - language models were stochastic in nature when that paper was written, they still are albeit we are trying to make them as deterministic as possible.

Re: What Emily Bender meant by "stochastic parrots"

#246

Earlier quoted context omitted.

I agree, which is why I think this species might be the start of something amazing: https://en.wikipedia.org/wiki/Larger_Pacific_striped_octopus

Is it a new species though? Or just a living fossil that exhibits behaviors that were bread out of all other octopi species? It's hard to imagine how high intelligence could be selected for without social component that provides the individuals a lot of pressure to outsmart each other.

Intelligence is plenty adaptive when you're just trying to outsmart the environment.

The notion that this is a "living fossil" is... pretty wild. It's genus octopus. Unless it's wildly misclassified (doubtful), it has a recent-ish common ancestor with most other common octopuses, which all have normal octopus behavior. You find yourself in need of extraordinary evidence.

Re: What Emily Bender meant by "stochastic parrots"

#247
post #197

Earlier quoted context omitted.

The authors were wrong about their core thesis and are now lying about it. That's the only criticism needed. They said, quote: > LMs are not performing natural language understanding (NLU), and only have success in tasks that can be approached by manipulating linguistic form ... which is presented as unarguable fact, yet is untrue. It was obviously wrong at the time it was written and it's been proven wrong in many w…

No, they were correct. In fact an LLM stitches together stuff it observed in its training data. That scales up way better than a lot of us expected, but it's still correct. If you train it on lots of working code, then it's useful for coding. If you trained it primarily on non-working code it would produce nonsense.

They do not "stitch together" anything. Neither on a technical level, nor a philosophical one. It "scales better than you expected" because your mental model is wrong.

And, not to insult you, but it's quite obviously wrong. As a mental model it fails to explain basic capabilities. How can an LLM follow elaborate instructions? How can it respond appropriately to user input, when the user input doesn't match any previously seen text? Hell - how does it even balance parentheses? There is no way to explain any of this without conceding that the LLM has semantic understanding. It knows that this comes after that, but "this" and "that" can be at an arbitrary level of abstraction.

Sure - they generate text "like" text they've seen before. That "like" does a ton of heavy lifting.

Re: What Emily Bender meant by "stochastic parrots"

#248

Earlier quoted context omitted.

A bird doesn't learn gravity or aerodynamics, it has no 'sense of physics'. It has sensory neural activity that a scientist can show is tied to these things, but you could, at least in theory, falsify the entire experience of the bird. There is nothing in a bird's brain that directly percieves reality. Colors and sounds and textures and so on are all false primitives that don't exist in nature without us, if that's c…

>A bird doesn't learn gravity or aerodynamics, it has no 'sense of physics'. That's not what I said. What I said was that it's physics that provides the ground truth. >you could, at least in theory, falsify the entire experience of the bird It wouldn't be a bird anymore, but a dysfunctional cyborg with false perceptions. >There is nothing in a bird's brain that directly percieves reality. Yes, of course there is. Ani…

>It wouldn't be a bird anymore, but a dysfunctional cyborg with false perceptions.

The criticism in the paper is of the architecture of LLMs, isn't it? The paper contends

"""Text generated by an LM is not grounded in communicative intent, any model of the world, or any model of the reader’s state of mind. [...] an LM is a system for haphazardly stitching together sequences of linguistic forms it has observed in its vast training data, according to probabilistic information about how they combine, but without any reference to meaning: a stochastic parrot"""

They're saying that the model cannot learn anything about reality irrespective of training data. Your point is an interesting one, but I think it's distinct. To your point though, this is unfalsifiable from the perspective of the "bird". I can't prove that I'm not a dysfunctional cyborg with false perceptions, which makes me wonder if that's a meaningful distinction.

>Animal sensory organs do not produce false information,

I happen to be in possession of some of these and I think this needs a "usually, under ordinary conditions". (Nitpicky and not critical to my point, but I liked the beginning of this sentence too much to edit it out)

>nor do they provide the brain an interpretation of what they perceive

Whereas this I'd argue is not true at all. My cones interact with wave-particle photons at particular wavelengths. I can't even conceptualize wave-particle duality (though some humans can), but "red" and "blue" are the bread and butter of my visual consciousness. These correspond to firings of my sensory neurons much more than they correspond to anything in reality. If that's not interpretation, what is it/what is interpretation?

>The word "real" is itself meaningless to it; they're both real in that they appear in its training corpus.

I expect we agree that I can show a multimodal model a real and a CGI picture and it can tell me which is which. I can take a CGI dragon to GPT 5 and say "look what I found in my backyard" and it will say "Yeah right". Are you saying this is something only possible thanks to RLHF or other modern techniques? That may be the case, unsure how to test that without access to pretraining-only models. Or would you say my experiment is faulty here and doesn't get to your underlying claim?

On the flipside, I could show some meh drawings of fairies to Arthur Conan Doyle, and he'd say "Whoa, this changes everything". I consider him to be one of the great rational minds of history, but he was unable to pass your test here. (In fairness, he was in his 60s and his senses may have dulled, though his belief in spiritualism at large dates to his prime).

Appreciate the conversation!

Re: What Emily Bender meant by "stochastic parrots"

#249

Earlier quoted context omitted.

Is it a new species though? Or just a living fossil that exhibits behaviors that were bread out of all other octopi species? It's hard to imagine how high intelligence could be selected for without social component that provides the individuals a lot of pressure to outsmart each other.

Intelligence is plenty adaptive when you're just trying to outsmart the environment. The notion that this is a "living fossil" is... pretty wild. It's genus octopus . Unless it's wildly misclassified (doubtful), it has a recent-ish common ancestor with most other common octopuses, which all have normal octopus behavior. You find yourself in need of extraordinary evidence.

> Intelligence is plenty adaptive when you're just trying to outsmart the environment.

I think there's very little evidence for that. Environment is just very slow when compared to outsmarting members of your own species that you tightly share the living space with. I don't think you appreciate how abnormal is high intelligence. How pathological conditions must have been to trigger development of something that bizarre and costly. This must have been run-away self-reinforcing feedback loop. We know of many such loops between predators and prey and none of them look strong and fast enough. Where we find intelligence, it's almost always tightly connected with social behaviors. It would be exceedingly strange if there was a true exception.

> The notion that this is a "living fossil" is... pretty wild. It's genus octopus.

I didn't mean that literally. What I meant was that the common ancestor of this octopus species and others might have been more like it, social, not like the rest. Maybe this species preserved of what normal octopus behavior used to be. If the solitary life became more evolutionarily favorable after the development of intelligence most descendant species could have switched.

> You find yourself in need of extraordinary evidence.

Yeah, maybe at some point we'll discover a huge collection of beaks in some under-ocean sediment and it will become obvious. Ocean floor is not really that explored.

Re: What Emily Bender meant by "stochastic parrots"

#250

Earlier quoted context omitted.

Intelligence is plenty adaptive when you're just trying to outsmart the environment. The notion that this is a "living fossil" is... pretty wild. It's genus octopus . Unless it's wildly misclassified (doubtful), it has a recent-ish common ancestor with most other common octopuses, which all have normal octopus behavior. You find yourself in need of extraordinary evidence.

> Intelligence is plenty adaptive when you're just trying to outsmart the environment. I think there's very little evidence for that. Environment is just very slow when compared to outsmarting members of your own species that you tightly share the living space with. I don't think you appreciate how abnormal is high intelligence. How pathological conditions must have been to trigger development of something that bizar…

> I think there's very little evidence for that.

The evidence is literally every other evolved form of intelligence. Including, despite your speculation, octopuses. Do recall "the environment" also includes other intelligent agents an animal needs to compete with. Does that help you understand? A runaway process might be required for human level intelligence but clearly not in general.

Basically all octopuses are solitary and die after breeding. LPSOs are nested in one group of octopuses. You're proposing that one group of octopuses developed social behavior and multiple breeding, then had a bunch of descendants who went back to exactly the normal octopus behavior. Hopefully it's obvious why this fails Occam's razor.

Post reply on HN