Live data from Hacker News

What Emily Bender meant by "stochastic parrots"

spectrum.ieee.org

211–220 of 280 posts

Re: What Emily Bender meant by "stochastic parrots"

#211
post #27

it annoys me how eager people are to hurl the word stochastic as pejorative. Statistics are a great tool for gleaning information from stochastic processes; statistics don't contribute randomness. Random sampling is necessary in order not to bias a sample, it's not used to contribute randomness to the sample but to preserve/measure the underlying distribution. (not meant to imply that training is random sampling)

It's a pejorative only because determinism is what makes computers useful in the first place. You get a consistent result, every single time, unlike if you have a human in the loop. Because LLMs are stochastic, they have removed the thing that makes computers useful to us, thus it's a pejorative.

What do you mean by determinism here? That you ask the computer 2+2 and it gives 4 as you expected or that if you ask the computer 2+2 and it hallucinates 5, you want it to always hallucinate 5?

Which one it doesn't do for you? Does it sometimes answer 4, sometimes 5?

There are definitely models that will always give 100% of the time the exact same answer, bit-for-bit, given the same input and seed. There are generative image models you can run locally doing just that. But you can also run some the SOTA chinese LLMs at "temperature 0" and, given the same input, they'll always give you the exact same output.

Because it's just a machine doing computation.

In the beginning of LLMs some "engineers" have tried to hand-wave non-sensical explanation as to why LLMs couldn't possibly be deterministic but: the open-weights models that can be run in a 100% deterministic way are way more powerful than the SOTA models of back then, so those explanation were pure rubbish bollocks.

Now of course if you run a complex chain of events, with LLMs doing calls to other LLMs, where some of them go fetch infos on unreliable networks, with infos that may have changed, then, logically, you won't always get the same answer.

Re: What Emily Bender meant by "stochastic parrots"

#212

Earlier quoted context omitted.

It's a pejorative only because determinism is what makes computers useful in the first place. You get a consistent result, every single time, unlike if you have a human in the loop. Because LLMs are stochastic, they have removed the thing that makes computers useful to us, thus it's a pejorative.

1. Determinism is a very small subset of what makes computers useful. Non-determinism like stochasticity is literally everywhere, like random seeds. 2. LLMs are detemrinistic. They have a parameter to tune how stochastic they are.

Sorry at first I downvoted you instead of upvoting you.

Re: What Emily Bender meant by "stochastic parrots"

#213

Earlier quoted context omitted.

If you want to see words form a shape I could point you towards concrete poetry, but I guess there is no point. Joyce wrote Finnegan’s Wake for 17 years and although superficially it seems complete gibberish, trodding through it you find meaning to words that are in no dictionary, sentence structures alien to English, etc. but still you are able to understand it, and perhaps some way the mind that produced it. So I d…

Oh, now I see where we have an actual difference of opinion. I don't think you can deny that even Finnegan's wake proceeds one token at a time; your interpretation of it may require more context or out-of-order interpretation, but that's just as true when observing text in German or Japanese, which have word ordering constraints that are alien to English speakers. How it was written is irrelevant; all we can observe…

I am no linguist, but I believe this is referred to as surface structure and deep structure. What you describe is a line of text that is grammatically somewhat adequate to pass as readable and you treat all text the same. when we read the text, we decipher meaning out of according to word references and syntax - just like with Python or C++. However, if this were solely the case, we could not read Finnegan’s Wake. Probably vast majority of modern poetry would be unread by anyone, as would be pretty much all major works of philosophy Kant onwards. Deep structure is what according to Chomsky et al. gives meaning to the language, ie. somewhat logical structure behind the mere words. English word strict has order, need it but actually not does. You skibidi rizz swag grok brah also, barely. We use language in the extended meaning of the word to create a model of the world and somehow the past riverrun skibidi transmits that model to others. This is what I believe Bender also tries to say in their paper.

Now, you could claim that LLM’s have this deep structure, create models of the world and are basically just like us, and certainly many here are adament that this is the case, being aghast how someone can “insult” LLM’s by calling them parrots. However, there really is not much proof to back up this belief. Usually LLM’s seem to copy existing surface structure from whatever source, and when it deviates from these patterns, it usually becomes incomprehensible. There’s much hoopla about LLM’s solving hard maths, but it seems even there they are mostly generalising from vast amounts of training data, rather than actually reasoning: https://arxiv.org/pdf/2410.05229

Re: What Emily Bender meant by "stochastic parrots"

#214

Earlier quoted context omitted.

To add to dwa3592's comment, a sentence is not a self-contained idea. The sentence doesn't include what any of the words in it mean, nor what "this sentence" refers to. The exact same sentence can mean different things depending on the text that surrounds it.

Fair point, and on its own it would be surprising to learn what "five" means from that sentence. But you can extrapolate- across a billion sentences, there will be "the next sentence has five words"s and "this sentence are grammared wrong" and so on. It would not be at all impossible to ground a world model on pure text for that reason. And 'not impossible' is sufficient to invalidate the paper's argument.

Let me provide a less superficial response, then.

>If multimodal models were still stochastic parrots by the original argument, humans would have to be as well; we don't have any way to ground anything beneath sense data

Animals don't passively learn from their perceptions, don't have a separation between training and inference, and don't have a prompt-response execution model. Besides its fundamental biology, the grounding an animal brain has is that when it outputs a motor signal it receives some feedback as to the effects of a signal of that strength. A bird learns to fly because the grounding truth of aerodynamics and gravity consistently respond in a specific way to the flapping of its wings. It doesn't learn by passively replaying thousands of hours of somatosensory recordings of flights.

A multimodal model doesn't have the capacity to do much with a prompt. It has no head to turn to look at an image from a slightly different angle to attempt to gleam more information, doesn't have the capacity to interact with the real thing the image represents in any way, and even if it requests another angle and is given it, it lacks the capacity to learn that new information permanently. A multimodal model knows about images of pipes and facts about pipes, but doesn't know pipes; it doesn't have literally first-hand experience with them.

>evolution can't have programmed some innate grounding into us because it didn't either.

What do you mean? Of course genetics programs ground truths. For example, "forward is the way your face points when your neck is relaxed" and "if you can feel it, then it's part of your body".

Re: What Emily Bender meant by "stochastic parrots"

#215

Earlier quoted context omitted.

Breaking it down this way is a great way to minimize the numbers so that it appears reasonable. See? Middle-Eastern investors are growing alfalfa in the western desert using legal allotments of water! That is so much worse than what we’re doing! Go after them! They can both be using an egregious amount of water for silly purposes. The other part of the water debate is also the pollution different systems create. Many…

some might call it "perspective"

Well the other thing is that agricultural uses are generally on a different scale and use different sources. And at least you can eat almonds and alfalfa. Probably more useful than asking an LLM to autocomplete your homework. Depending on your needs.

Even amongst municipal water usage, it seems hard to figure out how much DC usage factors in to other uses because often they’re not required, and therefore don’t, disclose their usage.

All I’m saying is that it’s not a closed case and not everyone complaining about DC water usage is a Russian propaganda bot.

Re: What Emily Bender meant by "stochastic parrots"

#216
post #200
post #196

Earlier quoted context omitted.

The thing that is off putting about how he uses rhetoric is that it feels like and-you deflection (tu quoque). > Claim that a data center is using 1000x as much water as a city of 88,000 people, where it’s actually using about 0.22x as much water as the city, and only 3% of the municipal water system the city relies on. She’s off by a factor of 4500. This is the single largest error in any popular book that I’ve foun…

> I think AI is powerful tool, but we still can't give DC expansion a pass. I really don't think we are. In the wider culture the idea that data centers use an absurd amount of water is baked in at this point. It frustrates me because I think it distracts from energy usage, which is a much more real issue than water usage.

I'd love to see DCs scaled with renewables, they would throttle down and run slower at night. The abstraction that power is available at whatever rate you want it is an expensive myth at scale.

I defer plugging in my electric car until 10am so I know that my neighbors solar arrays are charging my car.

In Seattle, every watt saved is water saved since we are blessed with hydro.

Re: What Emily Bender meant by "stochastic parrots"

#218
post #56

Earlier quoted context omitted.

It's such a tragedy that they're also extremely solitary animals and die shortly after reproducing the first (and only) time. Almost all other particularly intelligent animals seem to be gregarious, and it's easy to conclude that a social lifestyle tends to select for more intelligence, a sophisticated theory of mind, and so on (I like to think that that's exactly what was responsible for a runaway intelligence explo…

I agree, which is why I think this species might be the start of something amazing: https://en.wikipedia.org/wiki/Larger_Pacific_striped_octopus

Is it a new species though? Or just a living fossil that exhibits behaviors that were bread out of all other octopi species?

It's hard to imagine how high intelligence could be selected for without social component that provides the individuals a lot of pressure to outsmart each other.

Re: What Emily Bender meant by "stochastic parrots"

#219

Earlier quoted context omitted.

My criticism centers on the part of the paper they chose for their title, the “stochastic parrot” metaphor. And my criticism is that if you observe Claude code with opus 4.8 working through an entirely novel problem that nobody has ever worked on before and which certainly wasn’t in its training data, the choice to even metaphorically call them stochastic parrots turned out to be egregiously wrong. And secondarily, a…

oh boy! the lack of critical thinking here is staggering. >>And my criticism is that if you observe Claude code with opus 4.8 working through an entirely novel problem that nobody has ever worked on before and which certainly wasn’t in its training data, the choice to even metaphorically call them stochastic parrots turned out to be egregiously wrong. First of all, for your own benefit - Claude code stopped showing t…

I strongly suspect my critical thinking skills are better than yours.

For example, note my use of the phrase “turned out” and then consider whether your main point about the release date of plus 4.8 vs the paper makes any sense at all?

> how do you know what is novel?

Because I work on novel ASICs that were created by my company and have never been used or seen outside of my company?

Re: What Emily Bender meant by "stochastic parrots"

#220

Earlier quoted context omitted.

>>The contention that there is no grounding because the training data is linguistic and thus can only reference a world model is disproven in "This sentence has five words"- there's real, grounded information about what "five" means within that sentence. did you think this through? imagine the sentence was "This sentence has four words", now extrapolate that to all the shit that can exist in a dataset and train a mod…

I don't think this tone is at all justified. If you think otherwise, I do ask that you point out where I went too far in a comment that I feared was overburdened by caveats and admissions of my own human flaws. "This sentence has five words" is going to appear far more often than "This sentence has four words". This is the entire premise of LLMs working at all, stochastic parrots or otherwise.

you are right, i was more curt than i should have been. apologies. but you helped prove my point:

>>"This sentence has five words" is going to appear far more often than "This sentence has four words".

it's not about this at all. your point is about data quality. you need to take a step back. the point is that if you trained a language model just on this data set which has sentences akin to "this sentence has two words" - the model is going to learn that. this shows that the language modeling itself doesn't truly provide an understanding of the real world. you can train a language model with the most advanced technology on shitty data and the model will start providing shitty outputs confidently - the model will never say "hey there is something wrong with the data i am trained on". thats what 'understanding of the world' meant in that paper.

Post reply on HN