The author (bafflingly) seems to have completely missed the point- since anything they state up to page 15 (at which point I stopped reading) does not refute Chomsky's points at all. The author talks about LLMs and how they generate text and then goes on to talk about how it refutes Chomsky's claim about syntax and semantics. However it does not since Chomsky's primary claim is about how HUMANS acquire language. The…
Imagine being told all you need to do to learn Spanish, is to read a 300,000 word Spanish dictionary end to end so that you can probalistically come up with 1000 conversational phrases. Anyone who has learned a language can tell you it just doesn't work like that. You don't work by accumulating a massive dataset and training on it. No one can hold such a massive dataset of anything in their head at once.
Modern language models refute Chomsky’s approach to language
51–60 of 244 posts
Re: Modern language models refute Chomsky’s approach to language
#52Earlier quoted context omitted.
> The fact that you can replicate coherent text from probabilistic analysis and modeling of a very large corpus does not mean that humans acquire and generate language the same way. Also, the LLMs are cheating! They learned from us. It's entirely possible that you do need syntax/semantics/sapience to create the original corpus, but not to duplicate it. Let's see an AlphaZero-style version of an LLM, that learns langu…
>Also, the LLMs are cheating! No...they aren't. Humans aren't learning from thin air by any stretch of the imagination.
I don’t have a reference handy now (someone can probably do better) but I believe one way to see this is via the hearing impaired or hearing and sight impaired
Re: Modern language models refute Chomsky’s approach to language
#53Earlier quoted context omitted.
We did, at least once.
No, humans learn their languages through observation and generation - watching older people speak, and then imitating it. They get corrections, too, when they mispronounce or misuse words.
Re: Modern language models refute Chomsky’s approach to language
#54Earlier quoted context omitted.
Wordcels think LLMs imitate the human brain, when a shape rotator knows they really just imitate human language.
I love these new terms, can you elaborate on this?
Re: Modern language models refute Chomsky’s approach to language
#55Earlier quoted context omitted.
Wordcels think LLMs imitate the human brain, when a shape rotator knows they really just imitate human language.
Wordcels? Shape rotator?
> Wordcels are people who have high verbal intelligence and are good with words, but feel inadequately compensated for their skill. The term "cel" denotes frustration over being denied something they feel they deserve.1 Shape rotators are people with high visuospatial intelligence but low verbal intelligence, who have an intuition for technical problem-solving but are unable to account for themselves or apprehend historical context.2 The use of the terms has skyrocketed online in the past few months, especially in the last few days.0 The term "wordcel" is derived from incel and is used to describe someone who has high verbal intelligence but low "visuospatial" intelligence, whose facility for and love of complex abstraction leads them into rhetorical and political dead-ends.
Re: Modern language models refute Chomsky’s approach to language
#56Re: Modern language models refute Chomsky’s approach to language
#57Earlier quoted context omitted.
1. No Language model is yet even close to the scale of the human brain 2. Depending on what exactly you're trying to teach (perfect grammar, paragraphs of coherent text, basic reasoning), much less data is needed. https://arxiv.org/abs/2305.07759 3. Brains don't start at 0. Evolution, dna/rna etc. There's obviously some pre disposition for language learning in humans but that alone isn't enough ground for a "universa…
> No Language model is yet even close to the scale of the human brain If GPT-4 has 100 trillion parameters, it has as many parameters as the human brain has synapses. Synapses are a lot simpler than parameters; they're digital. A single neuron needs many synapses, all of roughly equal weight, emitting many pulses over a short time in order to convey a single weighted value. On top of that, you may have heard that the…
Re: Modern language models refute Chomsky’s approach to language
#58Earlier quoted context omitted.
Wordcels think LLMs imitate the human brain, when a shape rotator knows they really just imitate human language.
I love these new terms, can you elaborate on this?
Re: Modern language models refute Chomsky’s approach to language
#59We have known that it is possible to understand language without innateness. That is what linguists do.
If you look at how linguists know about innate features, the answer is almost always by first discovering them explicitly while analysing language data; not by opening a brain to see what is innately inside. [0]
The point about innateness is that it takes generations of linguists to learn from a blank slate properties of language that children learn in just years.
There are also numerous other arguments for innateness. From the way humans seem to spontaneously develop language in a language deprives environment, to the way language aquasition works being more consistent with other innate behaviours, to the pressence of weird properties that seem to be present across languages for no apparent logical reason.
The only insight I see from LLM is the same insight we have seen throughout macine learning. It is not nessasary to understand something if you can throw enough compute at it. This is powerful, and it enables us to do a lot, but it should not be confused with understanding.
[0] There are some instances leveraging MRI and other cognative research teqniques to get some insight into the inner workings of human language processing, but their role in developing current linguistics theory is thus far limited.
Re: Modern language models refute Chomsky’s approach to language
#60While I think some of the points in the article are interesting, the usual evidence for the Chomskian approach is the relative lack of input data for learning language by children in the wild. How much input data is used to train modern language models?
I agree. Question has not been settled. 20 year old human has * heard ~220 million words, talked 50 million words. * read ~10 million words. * experienced 420 million seconds of wakeful interaction with the environment (can be used to estimate the limit to conscious decisions, or number of distinct 'epochs' we experience) From a machine learning perspective human life is surprisingly small set of inputs and actions,…