Live data from Hacker News

Deciphering language processing in the human brain through LLM representations

research.google

91–100 of 111 posts

Re: Deciphering language processing in the human brain through LLM representations

#91

Earlier quoted context omitted.

As I said there are universal rules that human language processing follows (like hierarchical structure dependence); you can't have arbitrary syntax/grammars. It's true that science hasn't solved the main puzzles about how to characterize these rules. The fact that statistical models are better predictors than the-"true"-characterization-that-we-haven't-figured-out-yet is completely irrelevant, just as it would be ir…

> As I said there are universal rules that human language processing follows (like hierarchical structure dependence); you can't have arbitrary syntax/grammars. GP didn't say anything about grammars being arbitrary. In fact, his claim that grammars are models of languages would mean the complete opposite.

I don't think they have a consistent understanding of the word "grammar": they seem to use it in the grade-school sense (grammar for English, grammar for French) but then refer to Chomsky's universal grammar which is different (grammar rules that are common to all languages).

The main point of contention is their statement that "grammar follows language" which, in the Chomsky sense, is false: (universal) grammar/syntax describes the human language faculty (the internal language system) from which external languages (English, French, sign language) are derived, so (external) languages follow grammar.

Re: Deciphering language processing in the human brain through LLM representations

#92

Earlier quoted context omitted.

I just gave one: "Chris". Here's Chomsky describing the "Chris"-experiments ([1]) as part of a broader answer about how language is distinct from general cognition which I paraphrased above. > That doesn't contradict the argument that “thinking” and “language processing” are not two sequential or clearly separated modes in the brain but deeply intertwined. It's not an argument, it's an assertion, that is, in fact, co…

Like i said, these experiments stop at a vague 'Chris can still learn languages'. No comment on actual proficiency or testing. For all i know i can't have a meaningful conversation with this guy beyond syntactically correct speech. Or maybe the best proficiency he's ever managed is still pretty poor compared to the average human. I have no idea. There's no contradiction because i never argued/asserted the brain didn'…

It's irrelevant to the experiment: he could learn synthetic language with human-like grammar and could not learn synthetic languages with non-human-like grammar. Regular people could solve the non-human-like languages with difficulty. Because his language ability is much higher than his general problem solving ability it gives strong evidence that 1. human language capacity a special function, not a general purpose cognitive function and 2. it obeys a certain structure.

> There's no contradiction because i never argued/asserted the brain didn't have centers tuned for language, which is really all this experiment demonstrates.

I don't know what you are trying to say then.

Re: Deciphering language processing in the human brain through LLM representations

#93

Earlier quoted context omitted.

Like i said, these experiments stop at a vague 'Chris can still learn languages'. No comment on actual proficiency or testing. For all i know i can't have a meaningful conversation with this guy beyond syntactically correct speech. Or maybe the best proficiency he's ever managed is still pretty poor compared to the average human. I have no idea. There's no contradiction because i never argued/asserted the brain didn'…

It's irrelevant to the experiment: he could learn synthetic language with human-like grammar and could not learn synthetic languages with non-human-like grammar. Regular people could solve the non-human-like languages with difficulty. Because his language ability is much higher than his general problem solving ability it gives strong evidence that 1. human language capacity a special function, not a general purpose c…

>Because his language ability is much higher than his general problem solving ability

I don't see how you can say his language ability is much higher than his general problem solving ability if you don't know what proficiency of language he is capable of reaching.

When you are learning say English as a second language, there are proficiency tiers you get assigned when you get tested - A1, A2 etc

If he's learning all these languages but maxing out at A2 then his language ability is only slightly better than his general problem solving ability.

This is the point i'm trying to drive home. Maybe it's because i've been learning a second language for a couple years and so i see it more clearly but saying 'he learned x language' says absolutely nothing. People say that to mean anything from 'well he can ask for the toilet' to 'could be mistaken for a native'.

>I don't know what you are trying to say then.

The brain has over millions of years been tuned to speak languages with certain structures. Deviating from these structures is more taxing for the brain. True statement. But how on earth does that imply the brain isn't 'thinking' for the structures it is used to ? Do you say you did not think for question 1 just because question 2 was more difficult ?

Re: Deciphering language processing in the human brain through LLM representations

#94

Earlier quoted context omitted.

This is just wrong. Languages follow certain inviolable rules, most notably, hierarchical structure dependence. There are experiments (Moro, the subject "Chris") that show that humans don't process synthetic languages that violate these rules the same as synthetic languages that do (specifically it takes them longer to process and they use non-language parts of the brain to do so).

This does not mean that language in humans isn't probabilistic in nature. You seem to think that because there is structure then it must be rule based but that doesn't follow at all. When a group of birds fly, each bird discovers/knows that flying just a little behind another will reduce the amount of flaps it needs to fly. When you have nearly every bird doing this, the flock form an interesting shape. 'Birds fly in…

First, there is no evidence of any probabilistic processing at the level of syntax in humans (it's irrelevant what computers can do).

Second, I didn't say that, in language, structure implies deterministic rules, I said that there is a deterministic rule that involves the structure of a sentence. Specifically, sentences are interpreted according to their parse tree, not the linear order of words.

As for the birds analogy, the "rules" the birds follow actually does explain the V-shape that the flock forms. You make an observation "V-shaped flock" ask the question "why a V-shape and not some other shape" and try to find a explanation (the relative bird positions make it easier to fly [because of XYZ]). In the case of language you observe that there is structure dependence, you ask why it's that way and not another (like linear order) and try to come up with an explanation. You are trying to suggest that the observation that language has structure dependence is like seeing an image of an object in a cloud formation: an imagined mental projection that doesn't have any meaningful underlying explanation. You could make the same argument for pretty much anything (e.g. the double-slit experiment is just projecting some mental patterns onto random behavior) and I don't think it's a serious argument in this case either.

Re: Deciphering language processing in the human brain through LLM representations

#95

Earlier quoted context omitted.

This does not mean that language in humans isn't probabilistic in nature. You seem to think that because there is structure then it must be rule based but that doesn't follow at all. When a group of birds fly, each bird discovers/knows that flying just a little behind another will reduce the amount of flaps it needs to fly. When you have nearly every bird doing this, the flock form an interesting shape. 'Birds fly in…

First, there is no evidence of any probabilistic processing at the level of syntax in humans (it's irrelevant what computers can do). Second, I didn't say that, in language, structure implies deterministic rules, I said that there is a deterministic rule that involves the structure of a sentence. Specifically, sentences are interpreted according to their parse tree, not the linear order of words. As for the birds ana…

>First, there is no evidence of any probabilistic processing at the level of syntax in humans (it's irrelevant what computers can do).

There is plenty evidence for to suggest this

https://pubmed.ncbi.nlm.nih.gov/27135040/

https://pubmed.ncbi.nlm.nih.gov/25644408/

https://www.degruyter.com/document/doi/10.1515/9783110346916...

And research on syntactic surprisal—where more predictable syntactic structures are processed faster—shows a strong correlation between the probability of a syntactic continuation and reading times.

>In the case of language you observe that there is structure dependence, you ask why it's that way and not another (like linear order) and try to come up with an explanation. You are trying to suggest that the observation that language has structure dependence is like seeing an image of an object in a cloud formation: an imagined mental projection that doesn't have any meaningful underlying explanation.

No I'm suggesting that all you're doing here is cooking up some very nice fiction like Newton did when he proposed his model of gravity. Grammar does not even fit into rule based hierarchies all that well. That's why there are a million strange exceptions to almost every 'rule'. Exceptions that have no sensible explanations beyond, 'well this is just how it's used' because of course that's what happens when you try to break down an inherently probabilistic process into rigid rules.

Re: Deciphering language processing in the human brain through LLM representations

#96

Earlier quoted context omitted.

It's irrelevant to the experiment: he could learn synthetic language with human-like grammar and could not learn synthetic languages with non-human-like grammar. Regular people could solve the non-human-like languages with difficulty. Because his language ability is much higher than his general problem solving ability it gives strong evidence that 1. human language capacity a special function, not a general purpose c…

>Because his language ability is much higher than his general problem solving ability I don't see how you can say his language ability is much higher than his general problem solving ability if you don't know what proficiency of language he is capable of reaching. When you are learning say English as a second language, there are proficiency tiers you get assigned when you get tested - A1, A2 etc If he's learning all…

As I said it's not relevant but if you wanted to know you could put in the bare minimum of effort into doing your own research. From Smith and Tsimpli's "The Mind of a Savant": "On the [Gapadol Reading Comprehension Test] Christopher scored at the maximum level, indicating a reading comprehension of 16 years and 10 months". They describe the results of a bunch of other language tests, where he scores average to above average, including his translations of passages from a dozen different languages.

> But how on earth does that imply the brain isn't 'thinking' for the structures it is used to ? Do you say you did not think for question 1 just because question 2 was more difficult ?

The point isn't to define the word "thinking" it is to show that the language capacity is a distinct faculty from other cognitive capacities.

Re: Deciphering language processing in the human brain through LLM representations

#97

Earlier quoted context omitted.

First, there is no evidence of any probabilistic processing at the level of syntax in humans (it's irrelevant what computers can do). Second, I didn't say that, in language, structure implies deterministic rules, I said that there is a deterministic rule that involves the structure of a sentence. Specifically, sentences are interpreted according to their parse tree, not the linear order of words. As for the birds ana…

>First, there is no evidence of any probabilistic processing at the level of syntax in humans (it's irrelevant what computers can do). There is plenty evidence for to suggest this https://pubmed.ncbi.nlm.nih.gov/27135040/ https://pubmed.ncbi.nlm.nih.gov/25644408/ https://www.degruyter.com/document/doi/10.1515/9783110346916... And research on syntactic surprisal—where more predictable syntactic structures are processe…

> And research on syntactic surprisal—where more predictable syntactic structures are processed faster—shows a strong correlation between the probability of a syntactic continuation and reading times.

I'm not sure what this is supposed to show? If I can predict what you are going to say so what. I can predict you are going to pick something up too if you are looking at it and start moving your arm. So what?

The third paper looks like a similar argument. As far as I can tell neither paper 1 or 2 propose a probabilistic model for language. 1 talks about how certain language features are acquired faster with more exposure (that isn't inconsistent with a deterministic grammar). I believe 2 is the same.

> No I'm suggesting that all you're doing here is cooking up some very nice fiction like Newton did when he proposed his model of gravity.

Absolutely bonkers to describe Newton's model of gravity as "fiction". In that sense every scientific breakthrough is fiction: Bohr's model of the atom is fiction (because it didn't use quantum effects), Einstein's gravity will be fiction too when physics is unified with quantum gravity. No sane person uses the word "fiction" to describe any of this, it's just scientific refinement: we go from good models to better ones, patching up holes in our understanding, which is an unceasing process. It would be great if we could have a Newton-level "fictitious" breakthrough in language.

> Grammar does not even fit into rule based hierarchies all that well. That's why there are a million strange exceptions to almost every 'rule'. Exceptions that have no sensible explanations beyond, 'well this is just how it's used' because of course that's what happens when you try to break down an inherently probabilistic process into rigid rules.

No one is saying grammar has been solved, people are trying to figure out all the things that we don't understand.

Re: Deciphering language processing in the human brain through LLM representations

#98

Earlier quoted context omitted.

>First, there is no evidence of any probabilistic processing at the level of syntax in humans (it's irrelevant what computers can do). There is plenty evidence for to suggest this https://pubmed.ncbi.nlm.nih.gov/27135040/ https://pubmed.ncbi.nlm.nih.gov/25644408/ https://www.degruyter.com/document/doi/10.1515/9783110346916... And research on syntactic surprisal—where more predictable syntactic structures are processe…

> And research on syntactic surprisal—where more predictable syntactic structures are processed faster—shows a strong correlation between the probability of a syntactic continuation and reading times. I'm not sure what this is supposed to show? If I can predict what you are going to say so what. I can predict you are going to pick something up too if you are looking at it and start moving your arm. So what? The third…

>I'm not sure what this is supposed to show? If I can predict what you are going to say so what.

If the speed of your understanding varies with how frequent and predictable syntactic structures are then your understanding of syntax is a probabilistic process. A strictly non-probabilistic process would have a fixed, deterministic way of processing syntax, independent of how often a structure appears or how predictable it is.

>I can predict you are going to pick something up too if you are looking at it and start moving your arm. So what?

Ok ? This is very interesting. Do you seriously think this prediction right now isn't probabilistic ? You estimate not from rigid rules but past experience that it's likely I will pick it up. What if i push it off the table ? You think that isn't possible? What if i grab the knife in my bag while you're distracted and stab you instead? Probability is the reason you picked that option instead of the myriad of options.

>Absolutely bonkers to describe Newton's model of gravity as "fiction". In that sense every scientific breakthrough is fiction: Bohr's model of the atom is fiction (because it didn't use quantum effects), Einstein's gravity will be fiction too when physics is unified with quantum gravity. No sane person uses the word "fiction" to describe any of this, it's just scientific refinement: we go from good models to better ones, patching up holes in our understanding, which is an unceasing process. It would be great if we could have a Newton-level "fictitious" breakthrough in language.

"All models are wrong. Some are useful" - George Box. There's nothing insane with calling a spade a spade. It is fiction and many academics do view it in such a light. It's useful fiction, but fiction none the less. And yes, Einstein's theory is more useful fiction. Grammar is a model of language. It is not language.

Re: Deciphering language processing in the human brain through LLM representations

#99

Earlier quoted context omitted.

> And research on syntactic surprisal—where more predictable syntactic structures are processed faster—shows a strong correlation between the probability of a syntactic continuation and reading times. I'm not sure what this is supposed to show? If I can predict what you are going to say so what. I can predict you are going to pick something up too if you are looking at it and start moving your arm. So what? The third…

>I'm not sure what this is supposed to show? If I can predict what you are going to say so what. If the speed of your understanding varies with how frequent and predictable syntactic structures are then your understanding of syntax is a probabilistic process. A strictly non-probabilistic process would have a fixed, deterministic way of processing syntax, independent of how often a structure appears or how predictable…

> If the speed of your understanding varies with how frequent and predictable syntactic structures are then your understanding of syntax is a probabilistic process.

In what sense? I don't see how it tells you anything if you have the sentence "The cat ___ " and then you expect a verb like "went" but you could get a relative clause like "that caught the mouse". The sentence is interpreted deterministically not by what what follows after a fragment might contain but what it does contain. If you are more "surprised" by the latter it doesn't tell you that the process is not deterministic.

> Ok ? This is very interesting. Do you seriously think this prediction right now isn't probabilistic ? You estimate not from rigid rules but past experience that it's likely I will pick it up. What if i push it off the table ? You think that isn't possible. What if i grab the gun in my bag while you're distracted and shoot you instead?

I think you are confusing multiple things. I can predict actions and words, that doesn't mean sentence parsing/production is probabilistic (I'm not even sure exactly what a person might mean by that, especially with respect to production) nor does it mean arm movement is.

> "All models are wrong. Some are useful" - George Box. There's nothing insane with calling a spade a spade. It is fiction and many academics do view it in such a light. It's useful fiction, but fiction none the less. And yes, Einstein's theory is more useful fiction. Grammar is a model of language. It is not language.

I have no idea what you are saying: calling grammar a "fiction" was supposed to be a way to undermine it but now you are saying that it was some completely trivial statement that applies to the best science?

Re: Deciphering language processing in the human brain through LLM representations

#100

Earlier quoted context omitted.

This is just wrong. Languages follow certain inviolable rules, most notably, hierarchical structure dependence. There are experiments (Moro, the subject "Chris") that show that humans don't process synthetic languages that violate these rules the same as synthetic languages that do (specifically it takes them longer to process and they use non-language parts of the brain to do so).

Moro is apparently a reference to Andrea Moro, but I can't find any writing of his titled 'The Subject "Chris"'.

https://www.cambridge.org/core/books/abs/signs-of-a-savant/i...
Post reply on HN