Live data from Hacker News

What Emily Bender meant by "stochastic parrots"

spectrum.ieee.org

181–190 of 280 posts

Re: What Emily Bender meant by "stochastic parrots"

#181

Earlier quoted context omitted.

Yeah… LLMs clearly already have a world model

I think it's a good distinction to make between having and being, which seems to be what the whole "stochastic parrots" bit was intended to make all along. It doesn't make sense to say a model is in possession of its self. That's exactly the sort of poetic anthropomorphization that Bender was criticising here, and a good reason to not refer to an LLM as "an AI".

Wow this is exactly kind of pedanticism that annoys me with Bender. Glad that this kinda thing is becoming unpopular.

Re: What Emily Bender meant by "stochastic parrots"

#182
post #37

I paid a bit of attention to this paper and the phrase 'stochastic parrots' when it came out and i thought this was worth saying and doing at that time. their suggestions about financial and environmental costs are worth studying, their concern about carefully evaluating datasets to feed to the model rather than feeding the entire internet is fully justified. so - to everyone saying this was a bad paper; if you have…

The contention that there is no grounding because the training data is linguistic and thus can only reference a world model is disproven in "This sentence has five words"- there's real, grounded information about what "five" means within that sentence. While that's a trivial counterexample, I don't know that it's an obvious one (I didn't come up with it myself).

It's not a criticism of the paper itself, but multimodal models came shortly after and provide grounding that is more of the sort the paper is getting at, and it didn't seem like anybody updated on that at all. If multimodal models were still stochastic parrots by the original argument, humans would have to be as well; we don't have any way to ground anything beneath sense data and evolution can't have programmed some innate grounding into us because it didn't either. But (and maybe this is my own misperception) nobody threw in the towel at that point.

I confess I never read the original paper until now, opting to absorb by osmosis instead, and I was quite surprised that they don't really make a deeper case than that. After just a few paragraphs about how they can't be grounded because humans don't express their thoughts directly, it lurches into a page about how they can be biased by training. And they certainly can be, but that has little to say about their stochastic nature- humans are biased as a rule with no exception. (For the record, I only read the Stochastic Parrots section before this reply.)

It's not really a bad paper, but I don't see why it ever carried the esteem it did. Hating on it is like hating on Taylor Swift- she's fine, yes, but for her level of success, one is inclined to question every dumb lyric where others get a pass. (Apologies to Swift fans, substitute a successful artist you don't care for here.)

Re: What Emily Bender meant by "stochastic parrots"

#183

Earlier quoted context omitted.

I think it's a good distinction to make between having and being, which seems to be what the whole "stochastic parrots" bit was intended to make all along. It doesn't make sense to say a model is in possession of its self. That's exactly the sort of poetic anthropomorphization that Bender was criticising here, and a good reason to not refer to an LLM as "an AI".

Wow this is exactly kind of pedanticism that annoys me with Bender. Glad that this kinda thing is becoming unpopular.

There is more here then pedantry, but if you aren't interested in it, then go right ahead and live your life.

Re: What Emily Bender meant by "stochastic parrots"

#184

Bender's paper had this to say about stochastic parrots: "Contrary to how it may seem when we observe its output, an LM is a system for haphazardly stitching together sequences of linguistic forms it has observed in its vast training data, according to probabilistic information about how they combine, but without any reference to meaning: a stochastic parrot." This was not even a correct criticism in 2021. She is rig…

Even if pre-training was the only training step, it still wouldn't necessarily follow that the only thing the model is doing is stitching words together probabilistically, unless you expand the definition of "probabilistically" to the point that it becomes meaningless. This kind of thinking assumes that design of the training process and the "design" of the artifact that training produces must be similar.

Re: What Emily Bender meant by "stochastic parrots"

#185
post #175
post #135

Earlier quoted context omitted.

This is a useful piece on that: https://andymasley.substack.com/p/the-ai-water-issue-is-fake

How do you know this doesn't suffer from Gell-Mann Amnesia? The first version had so many glaring errors that have been "corrected" (removed), and I don't have the energy to comb through this one. I am highly skeptical of layperson debunking like this.

Andy has a good track record for writing about this. He shares plenty of credible citations - more so than most other people commenting in this space.

He also caught a major error in one of the most widely read books that helped kick off the whole data center water debate: https://blog.andymasley.com/p/empire-of-ai-is-wildly-mislead...

Re: What Emily Bender meant by "stochastic parrots"

#186
post #132
post #37

I paid a bit of attention to this paper and the phrase 'stochastic parrots' when it came out and i thought this was worth saying and doing at that time. their suggestions about financial and environmental costs are worth studying, their concern about carefully evaluating datasets to feed to the model rather than feeding the entire internet is fully justified. so - to everyone saying this was a bad paper; if you have…

My main criticism of the paper is that it says LLMs work "haphazardly", using probabilistic information. That is a hypothesis, but it is stated as a known fact, a fundamental limitation. It is true that LLMs often behave haphazardly, and do rely on statistics. But plenty of research has shown them behaving in methodical ways too. There are findings going both ways! Granted, many of the strongest contradictory results…

Statistical operation doesn’t preclude logical processing.

We’ve know that since 1943 when McCulloch-Pitts came up with the first “artificial neuron” definition. And since LLMs are a descendant technology — our assumption should be they’re reasoning in some internal learned logic.

This is what the evidence supports — eg, the “stochastic parrot” crowd never can explain transfer learning. Whereas for the internal reasoning crowd that is easy: removing your top level judgments from a theory still leaves you with useful terms for describing a new theory — eg, removing your judgments about “which animal is this?” but preserving the underlying structure for representing an image in your new judgments, “is this cancer?”

There’s 80 years of reason to think DNNs reason and zero support other than “sTaTs R mAgIc!” to support the stochastic parrot interpretation.

Ignorance isn’t argument.

Re: What Emily Bender meant by "stochastic parrots"

#187
post #48

> It argued that large language models (LLMs) generate text by statistically predicting likely sequences of words rather than understanding what they are saying—a process the authors captured with the metaphor of a “stochastic parrot,” a system that repeats patterns without comprehension. I don't understand what we're setting the record straight on. This is the core point of dispute, and the author just blazes past i…

I think it’s pretty clear that they are repeating without “comprehension” - both mechanistically (as in there is no facility for comprehension in their formulation) and in the ways they fail. The standard rs in strawberry, should I walk or drive to the car wash, etc all play on the fact that they don’t have any real world model or thoughts against which they can judge their output, as do many of the jailbreaks which…

> should I walk or drive to the car wash

there is an entire genre of riddles based on this kind of misdirection that works on humans, such as "A plane crashed on the border or US and Canada. Where do they bury the survivors?"

As with humans, comprehension is not a binary property of the agent - it is a quality that can be present in some situations and absent in others. LLMs may emit correct outputs sometimes because they do comprehend the input, and emit incorrect outputs in other cases when they do not comprehend the input.

In order to show that LLMs can't comprehend, we'd have to show that there are no (or at least very few) situations in which they exhibit comprehension, not show that there are some situations in which they don't.

Re: What Emily Bender meant by "stochastic parrots"

#188
post #37

I paid a bit of attention to this paper and the phrase 'stochastic parrots' when it came out and i thought this was worth saying and doing at that time. their suggestions about financial and environmental costs are worth studying, their concern about carefully evaluating datasets to feed to the model rather than feeding the entire internet is fully justified. so - to everyone saying this was a bad paper; if you have…

The authors were wrong about their core thesis and are now lying about it. That's the only criticism needed. They said, quote:

> LMs are not performing natural language understanding (NLU), and only have success in tasks that can be approached by manipulating linguistic form

... which is presented as unarguable fact, yet is untrue. It was obviously wrong at the time it was written and it's been proven wrong in many ways since. Worse is that they're still at it. In the article she's saying:

> Q: What are the most common misconceptions about the “stochastic parrots” metaphor? Bender: I think one of the biggest ones is, “Bender says AI is a stochastic parrot.”

Her name is on a paper titled "On the danger of stochastic parrots". It has a section titled "Stochastic parrots" and in section 6.1 it says:

> An LM is a system for haphazardly stitching together sequences of linguistic forms it has observed in its vast training data, according to probabilistic information about how they combine, but without any reference to meaning: a stochastic parrot.

She did say that, as clear as day. Now she's trying to rewrite history. Ugly behavior.

Re: What Emily Bender meant by "stochastic parrots"

#189

Earlier quoted context omitted.

I think part of it is she had excellent PR skills and a dedicated fan base. I was at Google when she quit, working in ML, and hadn't heard of her until the story broke. I remember there were a large number of Memegen posts about it, but no one I spoke with knew about her, so I assumed it was brigading. I think she's since since lost a lot of her allure, especially when she didn't change her mind when the facts about…

What exactly are the facts about AI water usage? I have trouble separating hysteria from reality but most of what I see still claims water usage is enormous

in my agi doomer opinion the water argument is akin to a concern that dropping a nuke may endanger certain rare species in the area

Re: What Emily Bender meant by "stochastic parrots"

#190

Earlier quoted context omitted.

The hysteria around water usage rests on people not knowing the scale of industrial civilization. First thing to do is compare any estimate of data center water usage with the water usage of almond farming. Or, if you want to focus on individual consumer choices, the water footprint of eating a hamburger.

Nice false dichotomy you got there there, might as well calculate the entire water usage for a a single GPU in the supply chain too. Something tells me one is extremely worse than the other when you account for all the water that's used in a single supply chain for high end electronics, but if you want to plop the measuring stick where ever along the whole pony show that makes you look better people will notice. Also…

let's compare ai water usage with something even more useless, water leaking out of pipes. the United States currently loses about 2 trillion gallons of water annually to leaking pipes. this is in comparison to around 230 billion gallons used by data centers.

we're literally dumping several times the amount of water used by data centers onto the ground for no benefit at all. oddly enough I haven't seen any protests about this despite how concerned everyone is about water usage.

Post reply on HN