Live data from Hacker News

How the Brain Parses Language

quantamagazine.org

81–90 of 90 posts

Re: How the Brain Parses Language

#81
post #6

> But what if our neurobiological reality includes a system that behaves something like an LLM? It almost seems like we got inspiration from our brain to build neural networks!

We've been making the same metaphor ("that's how the brain works") with each new major technology we come up with...

Re: How the Brain Parses Language

#82
post #6

> But what if our neurobiological reality includes a system that behaves something like an LLM? It almost seems like we got inspiration from our brain to build neural networks!

It isn’t clear though. Neural networks were inspired by the brain, but transformers? It is totally plausible but do we really think just in words?

>Neural networks were inspired by the brain, but transformers? It is totally plausible but do we really think just in words?

LLMs might be trained via words, but as a backend transformers are not just for words.

They're for high dimensional structured sequences. To make an analogy, transformers are not working on:

  Vector
but

  Vector
where words just happens to be a handy training set we use.

And, we too, might not think in words, but I bet that we do think using multi-dimensional sequences/vectors.

Re: How the Brain Parses Language

#83

Earlier quoted context omitted.

Yes, but at least now we're comparing artificial to real neural networks, so the way it works at least has a chance of being similar. I do think that a transformer, a somewhat generic hierarchical/parallel predictive architecture, learning from prediction failure, has to be at least somewhat similar to how we learn language, as opposed to a specialized Chompyskan "language organ". The main difference is perhaps that…

> comparing artificial to real neural networks I had a sad day in college when I thought I'd build my own ANN using C++. First thing I did was create a "Neuron" class, to mimic the idea of a human neuron. Second thing I did was realize that ANNs are actually just Weiner filters with a sigmoid on top. The base unit is not a "neuron".

Well, these exist tho:

https://en.wikipedia.org/wiki/Spiking_neural_network

Re: How the Brain Parses Language

#84
post #52

Earlier quoted context omitted.

>Yes, but at least now we're comparing artificial to real neural networks Given that the only similarity between the two of is just the "network" structure I'd say that point is pretty weak. The name "artificial neural network" it's just an historical artifact and an abstraction totally disconnected from the real thing.

Sure, but ANNs are at least connectionist, learning connections/strengths and representations, etc - close enough at that level of abstraction that I think ANNs can suggest how the brain may be learning certain things.

And these exist too:

https://en.wikipedia.org/wiki/Spiking_neural_network

Re: How the Brain Parses Language

#85
post #3

> It almost sounds like you’re saying there’s essentially an LLM inside everyone’s brain. Is that what you’re saying? >Pretty much. I think the language network is very similar in many ways to early LLMs, which learn the regularities of language and how words relate to each other. It’s not so hard to imagine, right? Yet, completely glosses over the role of rhythm in parsing language. LLMs aren’t rhythmic at all, are…

>Yet, completely glosses over the role of rhythm in parsing language.

If you're talking about speech cadence/rhythm, then we also parse written language which doesn't have that. And we're quite capable of parsing a monotone robotic voice speaking with a monotonous mechanical rhythm too.

Re: How the Brain Parses Language

#86

There's an interesting falsifiable prediction lurking here. If the language network is essentially a parser/decoder that exploits statistical regularities in language structure, then languages with richer morphological marking (more redundant grammatical signals) should be "easier" to parse — the structure is more explicitly marked in the signal itself. French has obligatory subject-verb agreement, gender marking on…

What do you make of this article? They used an auto-regressive genomic model to perform in-context learning experiments compared to language models. This showed that ICL behavior is not exclusive to language models. https://arxiv.org/html/2511.12797v1

Re: How the Brain Parses Language

#87

There's an interesting falsifiable prediction lurking here. If the language network is essentially a parser/decoder that exploits statistical regularities in language structure, then languages with richer morphological marking (more redundant grammatical signals) should be "easier" to parse — the structure is more explicitly marked in the signal itself. French has obligatory subject-verb agreement, gender marking on…

What do you make of this article? They used an auto-regressive genomic model to perform in-context learning experiments compared to language models. This showed that ICL behavior is not exclusive to language models. https://arxiv.org/html/2511.12797v1

This is great, thanks for the link. IMHO it actually supports the broader claim: if ICL emerges in both language models and genomic models, it suggests the phenomenon actually is about structure in the data, not something special about neural networks or transformers per se.

Genomes have statistical regularities (motifs, codon patterns, regulatory grammar). Language has statistical regularities (morphology, syntax, collocations). Both are sequences with latent structure. Similar architectures trained on either will repeat those structures.

That's consistent with my "instrumentation" view: the transformer is revealing structure that exists in the domain, whether that domain is English, French, or DNA. The architecture is the microscope; the structure was already there.

Re: How the Brain Parses Language

#88
post #77

Earlier quoted context omitted.

The confound concern is fair: no cross-linguistic comparison is perfectly controlled. The bet is that the effect size (if any) will be large enough to be informative despite the noise. But you're right that it's not ceteris paribus in a strict sense. Your proposal is interesting though. Synthetic manipulation of morphology within a single language. Have you seen this done? The challenge I'd anticipate is that "gender…

> The bet is that the effect size (if any) will be large enough to be informative despite the noise. But you have no grounds to ascribe it to the posited difference. Finding no effect might yield more information, but that's hard: given the amount of noise, you're bound to find a great many effects. > Have you seen this done? Not in LLMs, but there have been experiments with regularizing languages, and getting people…

On effect size: my primary goal at this stage is falsification. If French and English models show no meaningful differences at matched compute, that's informative: it would support the scaling hypothesis. If they do differ, I'll need to be careful about causal claims, but it would at least challenge the "transformers are magic" framing that treats architecture as the main story.

The L2 regularization and information theory pointers are helpful, it will go on my reading list. If you have favorites, I'll start there.

On the "we know nothing" point: I'm sympathetic. The Stroop example is exactly why I'm skeptical of strong claims in either direction. 197k papers and no mechanism suggests language processing has properties we don't yet have frameworks to describe. That's not mysticism. It's just acknowledging the gap between phenomenon and explanation.

Re: How the Brain Parses Language

#89
post #66

Earlier quoted context omitted.

I suspect you're more right than wrong. I'm a strong believer in this sort of thing -- that humans are best understood as a cyborg of a biological and semiotic organism, but mostly a "language symbiont inside a host". We should perhaps understand this as the strange creature of language jumping between hosts. But I suspect we're looking at a mule of sorts: it can't reproduce properly. But this mule could destroy us i…

Thank you for the Leiden references. I hadn't encountered this framework before. The "language symbiont" framing resonates with what I've been circling around: a system that operates with its own logic, sometimes orthogonal to conscious intention. The mule analogy is going to stick with me. LLMs have inherited the statistical structure of the symbiont without the host: pattern without grounding. Whether that makes th…

Glad I shared if it serves you!

> LLMs have inherited the statistical structure of the symbiont without the host: pattern without grounding.

I like this. I think it's not too far a leap to suggest something like "soul" without "body" -- a spirit in the truest sense. I think there's real value in the things we've believed ourselves to be made of though deep time, though without evidence or proper provenance. I suspect we've always been grappling to find language for the unnameable things.

Some of my own [somewhat outdated] reflections on language from the time I came across it, in case you're interested :) https://nodescription.net/notes/#2019-07-13

Re: How the Brain Parses Language

#90
post #51

Earlier quoted context omitted.

> language production and perception are quite separated in our heads Do you have any evidence for this? I am a former linguistics student (got my masters), and, after years of absenteeism in academia, interested in the current state of the affairs. So: "quite separated in our heads" Evidence for? against?

Afasia, and general measures of "normal" performance. There are various kinds of afasia, often linked to specific brain areas (Wernicke's and Broca's are well-known). And M/EEG and fMRI research suggests similar distinctions. It is difficult to reconcile with the idea that there is only one language system. And you will also have noticed that your skills in perception and production differ. You can read/listen better…

Ah, totally agreed. At least there is a clear auditory / motor part in the tasks that seems quite separate.

However, I find it also unlikely that the networks are totally separate, and I wonder if there are any evidence of areas that encode the "core/abstract" linguistic de/serialization (multidimensional and messy internal semantic information ←→ linear morphophonological information) both ways, or at least mechanism that manages to use gained input network competence to "train" or "manage" output network competence.

Why? Because even though, as you say, there is a differing performance in perception and production, there is also plenty of evidence of gaining linguistic competence from input, and then managing to convert that to performance in output.

Post reply on HN