Live data from Hacker News

What Emily Bender meant by "stochastic parrots"

spectrum.ieee.org

271–280 of 280 posts

Re: What Emily Bender meant by "stochastic parrots"

#271

Earlier quoted context omitted.

I am guessing (not asserting) that there is a sort of cap on water used for agriculture. It's possible we've already reached it. (?) So, on the matter of scale: there likely isn't a cap on water use of these datacenters. Both the heat emission and usage levels for these systems will likely go up unless there is a fundamental technical breakthrough. On the matter of utility: As a sibling of GP mentioned, the utility o…

what is the utility of alfaalfa? Because as it stands it's essentially just a way to export enormous amounts of water from a place where there's little (California) to a place where there's none (the Arabian peninsula). This is bullshit to be charitable.

Your alfalfa is a red herring. Agriculture encompasses more than a specific crop that we can always legislate against production for export. The data centers will not go away and the utility of the "AI" they are to enable is as of today dubious.

> This is bullshit to be charitable.

What is bullshit is sockpuppets shilling for the very few who have bet the 'farm' on this and want return on their investment, the wellbeing of society be damned.

Re: What Emily Bender meant by "stochastic parrots"

#272
post #268
post #260

Earlier quoted context omitted.

In fact LLMs are trained to predict the next token in the training set. Of course sometimes a new text input doesn't match the training set, or it matches two or more places in the training set. LLMs use a neural network to interpolate, so that's fine. Please look this up if you have any doubts. Ok. Now. I think you're adding something to the description above. Maybe what you're describing is something "emergent," or…

>LLMs use a neural network to interpolate "interpolate" in what vector space, pray tell? What does "interpolate" even mean, when I prompt it "write me a story about a sentient banana in the style of Hemingway and oh make it a commentary on class consciousness"? You can't assemble such a thing by cutting and pasting pieces of other text. That kind of "interpolation" has to happen at the semantic level - ipso facto, th…

>>Give any LLM the first paragraph of any Wikipedia article - almost certainly in the training set, and uniquely so - and it won't predict the next word correctly, a lot of the time. But it will predict a word that is grammatically correct, stylistically apropos, and most likely factually correct. So what's it really doing, hm?

you are grossly negligent of LLMs are created. I would highly recommend reading about post training, RLHF, alignment etc. First pass of training is literally "predict the next token". That's it. The first pass is also known as pre training. There's a shit ton of work (instructions, tool use, and reasoning etc) that's done afterwards because the pre-trained is model is useless. if you have some free time, I'd recommend doing this course - https://www.deeplearning.ai/courses/post-training-of-llms

Re: What Emily Bender meant by "stochastic parrots"

#273
post #266

Earlier quoted context omitted.

> More broadly, the predictive brain model would suggest that all of the brain, not just the neocortex, is dedicated to prediction. What would you say makes LLMs similar to the neocortex, rather than the basal ganglia or Broca's area? The whole brain is most definitely not dedicated to prediction, and I don't think "prediction" is a very useful model anyways. You could say the hippocampus is for "prediction" if you r…

Thank you for expanding! I come from more of an ML background so still learning a lot on these topics: Agreed that the neocortex uses fewer layers because of looping - I also suspect it's partly because neurons are more complex than the neurons used in ANNs, so in principle they should be capable of more sophisticated computation in a single forward pass (especially considering that handling multiple neurotransmitter…

I'm not sure if the complexity is at the neuron level. It's clearly possible; we see some pretty complex behavior from single celled organisms - Stentors - and there were the experiments (starting in the 70s, I believe) that conditioned flatworms, ground them up and fed them to new flatworms and showed that the new flatworms had the same conditioning. Those experiments were replicated but never explained, but they hint at something like maybe memories being encoded in RNA.

It'd be wild if our brains used a mechanism like that somewhere, but I doubt it's in the neocortex because those would be slow processes compared to what the neocortex has to do.

I think the sophistication in the neocortex just comes from it using something more sophisticated with transformers. Most of the looped LLM research I've scanned through, they seemed to be training the model to "know" when it should loop more and think harder, and I don't think that approach will ever work - you're solving for paths on a manifold, so the looping needs to happen at the abstraction layer below the manifold, looping more if it's still converging. And architectures with more recurrence are just harder for us humans to reason about and comparatively easier for dumb evolution to find clever solutions, so my money is that that's where the neocortex has LLMs beat.

And I'd say the manifold hypothesis is quite a bit more than a hypothesis at this point. There's been work on production LLMs to use manifold analysis to optimize them - there was a paper in Nature on it back in January. Baby stuff compared to what could be done, but the geometric approach keeps popping up in mechanistic interpretability and seems to be showing the most insights.

Re: What Emily Bender meant by "stochastic parrots"

#274

Earlier quoted context omitted.

> Intelligence is plenty adaptive when you're just trying to outsmart the environment. I think there's very little evidence for that. Environment is just very slow when compared to outsmarting members of your own species that you tightly share the living space with. I don't think you appreciate how abnormal is high intelligence. How pathological conditions must have been to trigger development of something that bizar…

> I think there's very little evidence for that. The evidence is literally every other evolved form of intelligence. Including, despite your speculation, octopuses. Do recall "the environment" also includes other intelligent agents an animal needs to compete with. Does that help you understand? A runaway process might be required for human level intelligence but clearly not in general. Basically all octopuses are sol…

> > I think there's very little evidence for that. > The evidence is literally every other evolved form of intelligence.

Not really, even harsh environments don't really select for intelligence. Neither does predator-prey dynamics. There are very few significantly intelligent species, and except for octopuses they are all social.

> Do recall "the environment" also includes other intelligent agents an animal needs to compete with.

If you include members of its own species in "the environment" we don't have any thing to discuss, because I fully agree that "the environment" understood as anything external to the individual can trigger the development of intelligence. Because it's trivial to observe that it did.

My claim is more informative (if true), that for the development of intelligence the environment must include populations on individuals living in close proximity, social, interacting a lot, where the strongest evolutionary pressure on a given species is the species itself. That doesn't seem to be true for most octopus species today. As an exception to the rule, their intelligence is interesting. I personally believe it comes from the same source as every other intelligence. And as the living space for octopuses expanded, they became more solitary. Maybe there's something general about this too. That with sufficiently intelligent species (perhaps if the intelligence can't rise anymore due to physical limitations), the other individuals present such a high danger that the species must spread apart and become more solitary. Maybe humans reached this stage already and spread apart, but as they ran out of Earth, they were forced back into old, more social mode, just now on the planetary scale. There is evidence that humans in Europe lived in way smaller groups and used to be way more murderous and cannibalistic not that long ago.

> Basically all octopuses are solitary and die after breeding.

That fits. There's no point in extending the life beyond reproduction if your kids are gona kill you anyways.

> You're proposing that one group of octopuses developed social behavior and multiple breeding, then had a bunch of descendants who went back to exactly the normal octopus behavior.

I'm not proposing that at all. What I'm proposing that all ancestors of octopuses were social and could breed multiple times. However, because of the proximity they became their greatest evolutionary pressure for themselves as they got more intelligent. Eventually they spread out due to some environmental change and their social behaviors and multiple breeding atrophied due to danger from other members of theirs own species, that now could just be nearly completely avoided. And that's the case for nearly all descendant species except this special one for some local environmental reason. This species shows us that social octopuses are (and were) a possibility.

What you consider normal octopus behavior is as normal as eyes of a molerat.

Re: What Emily Bender meant by "stochastic parrots"

#275

Earlier quoted context omitted.

> I think there's very little evidence for that. The evidence is literally every other evolved form of intelligence. Including, despite your speculation, octopuses. Do recall "the environment" also includes other intelligent agents an animal needs to compete with. Does that help you understand? A runaway process might be required for human level intelligence but clearly not in general. Basically all octopuses are sol…

> > I think there's very little evidence for that. > The evidence is literally every other evolved form of intelligence. Not really, even harsh environments don't really select for intelligence. Neither does predator-prey dynamics. There are very few significantly intelligent species, and except for octopuses they are all social. > Do recall "the environment" also includes other intelligent agents an animal needs to…

That's a fun story, but it doesn't match the evidence at any point. Your fixation on intraspecific competition as the only significant driver of intelligence is leading you far astray.

Re: What Emily Bender meant by "stochastic parrots"

#276

Earlier quoted context omitted.

> > I think there's very little evidence for that. > The evidence is literally every other evolved form of intelligence. Not really, even harsh environments don't really select for intelligence. Neither does predator-prey dynamics. There are very few significantly intelligent species, and except for octopuses they are all social. > Do recall "the environment" also includes other intelligent agents an animal needs to…

That's a fun story, but it doesn't match the evidence at any point. Your fixation on intraspecific competition as the only significant driver of intelligence is leading you far astray.

The evidence is the lack of any solitary, highly intelligent species apart from octopuses.

Re: What Emily Bender meant by "stochastic parrots"

#277
post #260

Earlier quoted context omitted.

It requires some level of semantic understanding, like what a paren is and what it means to balance them. The issue in this discussion is that "predict the next token" is a problematically reductive description of what's going on. It's like saying compilers are programs that emit bytes or that humans are mammals that make sounds. It's not strictly false but it's not capturing the depth of what's happening either. A s…

In fact LLMs are trained to predict the next token in the training set. Of course sometimes a new text input doesn't match the training set, or it matches two or more places in the training set. LLMs use a neural network to interpolate, so that's fine. Please look this up if you have any doubts. Ok. Now. I think you're adding something to the description above. Maybe what you're describing is something "emergent," or…

Ah, OK. I called her a liar not because of a dispute over what exactly transformers are doing inside, but because in TFA we see:

> What are the most common misconceptions about the “stochastic parrots” metaphor?

Bender: I think one of the biggest ones is, “Bender says AI is a stochastic parrot.”

But in the paper itself we see her say exactly that, several times. What she's trying to do now is wordsmith out of it by claiming that in her world LLMs are totally unrelated to AI, so when she said LLMs are stochastic parrots she wasn't making a claim about AI.

Nobody else defines AI to exclude LLMs, nor did they at the time, and it wasn't the core of the argument she made either. But then she admits that the paper does "generalize towards AI" at the end. So... whatever.

It's morally important to reject this stuff. When academics play word games it devalues all institutional output.

Re: What Emily Bender meant by "stochastic parrots"

#278
post #260

Earlier quoted context omitted.

In fact LLMs are trained to predict the next token in the training set. Of course sometimes a new text input doesn't match the training set, or it matches two or more places in the training set. LLMs use a neural network to interpolate, so that's fine. Please look this up if you have any doubts. Ok. Now. I think you're adding something to the description above. Maybe what you're describing is something "emergent," or…

Ah, OK. I called her a liar not because of a dispute over what exactly transformers are doing inside, but because in TFA we see: > What are the most common misconceptions about the “stochastic parrots” metaphor? Bender: I think one of the biggest ones is, “Bender says AI is a stochastic parrot.” But in the paper itself we see her say exactly that, several times. What she's trying to do now is wordsmith out of it by c…

I see. Ok.

Re: What Emily Bender meant by "stochastic parrots"

#280
post #233

Earlier quoted context omitted.

> Or, if you want to focus on individual consumer choices, the water footprint of eating a hamburger. To drive this point home: if every American ate exactly one less hamburger per year, it would entirely offset the annual water consumption of all US datacenters (including, therefore, the water footprint of AI).

> the annual water consumption of all US datacenters How many hamburgers will it be if all the proposed/pending datacenters are completed and running?

I would expect within the same order of magnitude. Worst-case, skipping a burger once a month instead of once a year is hardly some insurmountable sacrifice (and I say that as a fatass who loves himself a good burger).
Post reply on HN