Live data from Hacker News

Stochastic Parrots: Frequently Unasked Questions

medium.com

41–50 of 64 posts

Re: Stochastic Parrots: Frequently Unasked Questions

#41

> Most things we historically do with computing are not well approximated by extruding synthetic text. I don't understand this point. I feel like almost everything associated with computing is extruding synthetic text.

It seems like a criticism that's actually a hint at a bigger point. The entire appeal/hype is due to the promise of doing things that historically computers have not done well.

That's captured elsewhere - attempts to create "synthetic human behavior" - but mostly around ethics vs practical function or consumer appeal.

Even just a "stochastic parrot" can be extremely valuable if the parrot is fast enough and can connect enough dots in a human-reasoning-style to say things like "what could come after a description of a problem, some background info, and a question about what could have caused the problem? Probably a relevant hypothesis that fits the background facts and the problem description" and then generate a high-probability-fitting sequence of text to spit out.

There doesn't need to be any more intent in that than just "predict what would be the next text that would be similarly connected to the previous in the same way text in the model training process would." It doesn't need to be intending to solve the problem if the hit rate is good enough such that predicting how someone else would describe the solution is often the same as actually "intending" to solve it...

Nor does the ability to predict things stochasticly mean that there isn't any symbolic way to do the same. Quite possibly the stochastic process is just a brute-force rough approximation of what a true symbolic model could do. IMO the success of the stochastic approach is exactly in line with the existence of some sort of underlying structure/system. (Though such as system would have to be incredibly complex to support all the crazy things we do with language.)

Re: Stochastic Parrots: Frequently Unasked Questions

#42

It would have been nice to see some version of “I am very surprised by how far LLMs have come since I wrote the stochastic parrots paper, here is how I have revised my thinking.” But there is nothing like that and the author is just doubling down or trying to correct perceived “misinterpretations” of her work. Meanwhile you have multiple Fields Medalists (Tau, Gowers) saying they’re very impressed by LLMs’ mathematic…

It's clear from this comment that you did not read the full article. If you did then you'd have seen that the author addresses this criticism you're making here.

Re: Stochastic Parrots: Frequently Unasked Questions

#43

It would have been nice to see some version of “I am very surprised by how far LLMs have come since I wrote the stochastic parrots paper, here is how I have revised my thinking.” But there is nothing like that and the author is just doubling down or trying to correct perceived “misinterpretations” of her work. Meanwhile you have multiple Fields Medalists (Tau, Gowers) saying they’re very impressed by LLMs’ mathematic…

The Parrots paper: "Contrary to how it may seem when we observe its output, an LM is a system for haphazardly stitching together sequences of linguistic forms it has observed in its vast training data, according to probabilistic information about how they combine, but without any reference to meaning: a stochastic parrot." So perhaps this has always been a negative claim, about what language model AI is not .

> "Contrary to how it may seem when we observe its output, an LM is a system for haphazardly stitching together sequences of linguistic forms it has observed in its vast training data, according to probabilistic information about how they combine, but without any reference to meaning: a stochastic parrot."

and

> "Meanwhile you have multiple Fields Medalists (Tau, Gowers) saying they’re very impressed by LLMs’ mathematical reasoning, something that the stochastic parrots thesis (if it has any empirically-predictive content at all) would predict was impossible. I doubt Tau and Gowers thought much of LLMs a few years ago either. But they changed their minds. Who do you want to listen to?"

I don't understand how these things are supposedly incompatible.

Larger models and further other refinement reduce the "haphazardness" of produced text. A big enough model with enough semantic connections between different words/phrasings/etc plus enough logical connections of how cause and effect, question and answer, works in human language can obviously stitch together novel sequences when presented with novel prompts. (The output was not limited to sequences of n words that appeared 1:1 in the training data for any n for at least three and a half years now, if not even back to when the paper was written.)

"without any reference to meaning" veers into the philosophical (see how much "intent" is brought up in the linked post today). But has anything been proven wrong about the idea that the text prediction is based on probabilistic evaluation based on a model's training data? E.g. how can you prove "reasoning" vs "stochastic simulated reasoning" here?

Perhaps a useful counterfactual (but hopelessly-expensive/possibly-infeasible) would be to see if you could program a completely irrational LLM. Would such a model be able to "reason" it's way into realizing its entire training model was based on fallacies and intentionally-misleading statements and connections, or would it produce consistent-with-its-training-but-logically-wrong rebuttals to attempts to "teach" it the truth?

Re: Stochastic Parrots: Frequently Unasked Questions

#44
post #6

Earlier quoted context omitted.

TIL that the "merest resemblance of thinking" is enough to take gold at IMO.

Automated theorem provers are not new, in fact they are very old. One of the most automated is ACL2, which uses the well studied waterfall method (unrelated to waterfall development). LLMs certainly use something similar, except they understand text as input. LLMs, especially used for marketing stunts, have way more computing power available than any theorem prover ever had. They probably do random restarts if a proo…

LLMs certainly use something similar

They certainly do not. Read the papers where the IMO results were presented. No tools of any kind were used.

Re: Stochastic Parrots: Frequently Unasked Questions

#45

> Most things we historically do with computing are not well approximated by extruding synthetic text. I don't understand this point. I feel like almost everything associated with computing is extruding synthetic text.

Just to name some of the main things I think of computers doing, especially with a historical lens: analyzing data, processing transactions, simulating dynamics of physical systems, controlling electronic parts of devices, providing entertainment, encoding/decoding audio/video/text. I think these are the kinds of things that Dr Bender is saying are not well suited to textual tools.

Re: Stochastic Parrots: Frequently Unasked Questions

#46
post #13

Earlier quoted context omitted.

No? There's no model involved. It's all just probabilistic. LLMs understand what you're thinking as well as a mood ring.

Nothing about an LLM is “just”. In what precise sense do you mean it is probabilistic?

There's a reason stochastic was used in the original phrase instead of "probabilistic."

While most inference executions are intentionally non-deterministic, even a purely deterministic one would still be stochastic in that the model itself was built in a process such that the statistical frequency, sequencing, etc of the training text and followup processes all heavily influence the result.

Because of that, the output is the sort of thing that is not expected to generate 100% perfect output 100% of the time, but to have a good probability of being like-in-kind-to-the-training-data (and useful/relevant as a result).

(As compared to a non-stochastic model, like arithmetic on integers, where 2+2 is always gonna be 4 and you don't have a chance of coming up with some novel pair of inputs to addition that will cause your arithmetic to miss the mark.)

Re: Stochastic Parrots: Frequently Unasked Questions

#47
post #2

Lovely article well worth attention by virtue of its regard for the cultural traits of terminology and its inflections, while also debunking the pervasive lore that "AI" devices are doing anything but the merest resemblance of thinking. It's rare to read an author who can directly face Brandolini's Law of misinformation asymmetry and not only hold his own against the bullshit but overcome it.

You're dismissing LLM-generated text as the "merest resemblance of thinking" when the way it resembles thinking is becoming increasingly useful.

When I prompt a coding agent to fix a bug, it outputs text describing a hypothesis and more text that results in running shell commands to test the hypothesis. If the output shows that it guessed wrong, it outputs more text to test a different hypothesis, and more text to edit code, and in the end, the bug is fixed.

The text resembles the output of a reasoning process closely enough to actually work. Maybe, for some purposes, it doesn't matter if it's "real" or not?

What does "real" reasoning do for us that the imitation doesn't do? Does it come up with better hypotheses? Is it better at testing them? Sometimes, but not always. Human reasoning is more expensive, less available, and sometimes gets poor results.

Re: Stochastic Parrots: Frequently Unasked Questions

#48

"Text generated by an LM is not grounded in communicative intent, any model of the world, or any model of the reader’s state of mind." Modelling text describing the world is not modelling (some aspect) of the world? Modelling the probability that a reader likes or dislike a piece of text is not modelling (some aspect) of a reader's state of mind?

>Modelling text describing the world is not modelling (some aspect) of the world?

The text describes the world to humans. This is the crucial thing that you miss. It is very subjective.

Imagine that you learn the grammar of a foreign language without learning the meaning of the words. You might be able to make grammatically valid sentences. But you will still will not understand a single thing that something written in that language describes. But that will be perfectly clear to someone who actually understand the meaning of the words.

When you train LLMs on large volumes of text that describe logically consistent facts in a million different ways, the "logic" sort of becomes part of the grammer that the model learns. That is logic becomes a higher kind of "grammer" or a enormous set of grammatical rules that it captures. But that does not mean the model can do actual logic.

Re: Stochastic Parrots: Frequently Unasked Questions

#49
post #27

Earlier quoted context omitted.

There is also no evidence that Radio Free Europe is still linked to the CIA. Just look at the donors of Renaissance Misanthropy. But we are feeding a sealion who does not know how the math proof logic in LLMs work, probably because it is a highly computationally expensive random restart hack calling Lean that is unpublishable.

Many of these results don't rely on repeatedly calling Lean. You have no clue what you're talking about. > Just look at the donors of Renaissance Misanthropy. If you're actually interested, who funds each project is listed in the PDF here. https://www.renaissancephilanthropy.org/annual-reports As you can see, it's mainly philanthropic projects of wealthy families.

They literally operate on the model developed by Kleiner Perkins:

https://www.renaissancephilanthropy.org/the-fund-model

But get your AI friends to downvote truth and sink the entire submission, because that is how the AI fascists operate.

Re: Stochastic Parrots: Frequently Unasked Questions

#50

It would have been nice to see some version of “I am very surprised by how far LLMs have come since I wrote the stochastic parrots paper, here is how I have revised my thinking.” But there is nothing like that and the author is just doubling down or trying to correct perceived “misinterpretations” of her work. Meanwhile you have multiple Fields Medalists (Tau, Gowers) saying they’re very impressed by LLMs’ mathematic…

It's clear from this comment that you did not read the full article. If you did then you'd have seen that the author addresses this criticism you're making here.

I did read it. She doesn’t mention mathematics or RLVR training once, so I assume you’re referring to my point about empirical testability. Well, I think her statement that the claim “LLMs are stochastic parrots” is not an empirical claim is false, and she’s being disingenuous there with a classic motte-and-bailey fallacy. She quotes her own original paper thus:

> Text generated by an LM is not grounded in communicative intent, any model of the world, or any model of the reader’s state of mind. It can’t have been, because the training data never included sharing thoughts with a listener, nor does the machine have the ability to do that. This can seem counter-intuitive given the increasingly fluent qualities of automatically generated text, but we have to account for the fact that our perception of natural language text, regardless of how it was generated, is mediated by our own linguistic competence and our predisposition to interpret communicative acts as conveying coherent meaning and intent, whether or not they do [89, 140]. The problem is, if one side of the communication does not have meaning, then the comprehension of the implicit meaning is an illusion arising from our singular human understanding of language (independent of the model). Contrary to how it may seem when we observe its output, an LM is a system for haphazardly stitching together sequences of linguistic forms it has observed in its vast training data, according to probabilistic information about how they combine, but without any reference to meaning: a stochastic parrot.

Do you really think that claiming the output of an LLM “has no reference to meaning” is not an empirical claim? That it doesn’t attempt to place any bounds whatsoever on what LLMs can and cannot do? LLMs can solve some very difficult mathematical problems quite well now: see the article from Gowers that was on here recently. Do you think that the output in a situation like that “has no reference to meaning?” If so, you’ll have to explain why, because I don’t understand at all.

Post reply on HN