Live data from Hacker News

Stochastic Parrots: Frequently Unasked Questions

medium.com

51–60 of 64 posts

Re: Stochastic Parrots: Frequently Unasked Questions

#51
post #48

"Text generated by an LM is not grounded in communicative intent, any model of the world, or any model of the reader’s state of mind." Modelling text describing the world is not modelling (some aspect) of the world? Modelling the probability that a reader likes or dislike a piece of text is not modelling (some aspect) of a reader's state of mind?

>Modelling text describing the world is not modelling (some aspect) of the world? The text describes the world to humans. This is the crucial thing that you miss. It is very subjective. Imagine that you learn the grammar of a foreign language without learning the meaning of the words. You might be able to make grammatically valid sentences. But you will still will not understand a single thing that something written…

Thanks for your explanation, I find it much more intuitive than the paper's.

In your opinion, does a Calculus solver model certain aspects of the world?

Re: Stochastic Parrots: Frequently Unasked Questions

#52
post #48

"Text generated by an LM is not grounded in communicative intent, any model of the world, or any model of the reader’s state of mind." Modelling text describing the world is not modelling (some aspect) of the world? Modelling the probability that a reader likes or dislike a piece of text is not modelling (some aspect) of a reader's state of mind?

>Modelling text describing the world is not modelling (some aspect) of the world? The text describes the world to humans. This is the crucial thing that you miss. It is very subjective. Imagine that you learn the grammar of a foreign language without learning the meaning of the words. You might be able to make grammatically valid sentences. But you will still will not understand a single thing that something written…

> When you train LLMs on large volumes of text that describe logically consistent facts in a million different ways, the "logic" sort of becomes part of the grammer that the model learns. That is logic becomes a higher kind of "grammer" or a enormous set of grammatical rules that it captures. But that does not mean the model can do actual logic.

This is the kind of stuff people were saying in 2023. But it’s 2026 now and LLMs aren’t just trained by reading lots of text anymore. That’s “pretraining”, and it’s still the first stage, but LLMs also have a huge amount of RLVR training where they actually do solve huge numbers of mathematical and logic puzzles and update their weights in response. They don’t just learn mathematics from reading about it now. They learn it by doing it. That is why they can now solve hard problems and probe theorems.

> that does not mean the model can do actual logic.

But they do, all the time. (Please tell me you’ve at least put a frontier LLM through its paces in the last 6 months?) If you think they can’t do logic and reasoning, can you provide examples of specific math or logic problems that you think a frontier LLM can’t do?

Re: Stochastic Parrots: Frequently Unasked Questions

#53
post #48

Earlier quoted context omitted.

>Modelling text describing the world is not modelling (some aspect) of the world? The text describes the world to humans. This is the crucial thing that you miss. It is very subjective. Imagine that you learn the grammar of a foreign language without learning the meaning of the words. You might be able to make grammatically valid sentences. But you will still will not understand a single thing that something written…

> When you train LLMs on large volumes of text that describe logically consistent facts in a million different ways, the "logic" sort of becomes part of the grammer that the model learns. That is logic becomes a higher kind of "grammer" or a enormous set of grammatical rules that it captures. But that does not mean the model can do actual logic. This is the kind of stuff people were saying in 2023. But it’s 2026 now…

>If you think they can’t do logic and reasoning, can you provide examples of specific math or logic problems that you think a frontier LLM can’t do?

When a thing can "solve" a complex math problem without having the ability to count, then it is clear that this things is not "reasoning" and doing "logic".

Re: Stochastic Parrots: Frequently Unasked Questions

#54
post #48

"Text generated by an LM is not grounded in communicative intent, any model of the world, or any model of the reader’s state of mind." Modelling text describing the world is not modelling (some aspect) of the world? Modelling the probability that a reader likes or dislike a piece of text is not modelling (some aspect) of a reader's state of mind?

>Modelling text describing the world is not modelling (some aspect) of the world? The text describes the world to humans. This is the crucial thing that you miss. It is very subjective. Imagine that you learn the grammar of a foreign language without learning the meaning of the words. You might be able to make grammatically valid sentences. But you will still will not understand a single thing that something written…

> Imagine that you learn the grammar of a foreign language without learning the meaning of the words. You might be able to make grammatically valid sentences. But you will still will not understand a single thing that something written in that language describes. But that will be perfectly clear to someone who actually understand the meaning of the words.

so... back to chinese room arguments?

just because amazon worker inside is just moving folders around following rules, doesn't by default mean the room as a whole can't be corresponding to "something that doesn't understand"

denying emergence as a phenomenon isn't useful when "there are plenty of higher abstraction levels in multiple fields that still capture 99% of events and are easier to model and react to" is the counterpoint

Re: Stochastic Parrots: Frequently Unasked Questions

#55
I found it ironic that in an article by a brilliant linguist (and I mean that genuinely) about how fluidity of language can fool us into perceiving logic that isn't there that I was thrown out of frame by a silly grammatical error that an LLM would never make:

"The more direct inspiration for me was an email from Stuart Russell in September 2020 to Alexander Koller and I about our ACL 2020 paper".

It's a good article, and worth reading. But somehow I also have enough trouble understanding how she would make this error that---in the inverse of the fluidity argument---I start to doubt her the rest of her logic based on one silly irrelevant grammar mistake.

Re: Stochastic Parrots: Frequently Unasked Questions

#56

It would have been nice to see some version of “I am very surprised by how far LLMs have come since I wrote the stochastic parrots paper, here is how I have revised my thinking.” But there is nothing like that and the author is just doubling down or trying to correct perceived “misinterpretations” of her work. Meanwhile you have multiple Fields Medalists (Tau, Gowers) saying they’re very impressed by LLMs’ mathematic…

...did you read TFA?

The Hacker News guidelines say “Please don't comment on whether someone read an article. ‘Did you even read the article? It mentions that’ can be shortened to ‘The article mentions that’.”

https://news.ycombinator.com/newsguidelines.html

Re: Stochastic Parrots: Frequently Unasked Questions

#57
post #13

Earlier quoted context omitted.

No? There's no model involved. It's all just probabilistic. LLMs understand what you're thinking as well as a mood ring.

The model is the thing which is learned in order to make the probabilistic prediction with low entropy.

Well this is probably the same kind of semantic trap she's fighting with. Yes, you're right it's a model. The distinction is that they models of _language_ and not thoughts or feelings.

Re: Stochastic Parrots: Frequently Unasked Questions

#58
post #53

Earlier quoted context omitted.

> When you train LLMs on large volumes of text that describe logically consistent facts in a million different ways, the "logic" sort of becomes part of the grammer that the model learns. That is logic becomes a higher kind of "grammer" or a enormous set of grammatical rules that it captures. But that does not mean the model can do actual logic. This is the kind of stuff people were saying in 2023. But it’s 2026 now…

>If you think they can’t do logic and reasoning, can you provide examples of specific math or logic problems that you think a frontier LLM can’t do? When a thing can "solve" a complex math problem without having the ability to count, then it is clear that this things is not "reasoning" and doing "logic".

You didn’t answer my question. You just restated your claims.

Specific examples? Specific tasks?

Re: Stochastic Parrots: Frequently Unasked Questions

#59

Earlier quoted context omitted.

It's clear from this comment that you did not read the full article. If you did then you'd have seen that the author addresses this criticism you're making here.

I did read it. She doesn’t mention mathematics or RLVR training once, so I assume you’re referring to my point about empirical testability. Well, I think her statement that the claim “LLMs are stochastic parrots” is not an empirical claim is false, and she’s being disingenuous there with a classic motte-and-bailey fallacy. She quotes her own original paper thus: > Text generated by an LM is not grounded in communicat…

The fact that you disagree with a claim does not automatically make it an empirical claim.

Re: Stochastic Parrots: Frequently Unasked Questions

#60
post #57

Earlier quoted context omitted.

The model is the thing which is learned in order to make the probabilistic prediction with low entropy.

Well this is probably the same kind of semantic trap she's fighting with. Yes, you're right it's a model. The distinction is that they models of _language_ and not thoughts or feelings.

When I read your reply, I’m also modeling language. Tokens are just the discretization of the model’s eyes and ears. My brain does a huge amount of work to represent what’s happening in the world based on discrete information received from the outside world, just like language models do.
Post reply on HN