Also, dogs and cats, and even our close relatives, primates, don't develop a capacity to language.
Yes, and that's one of my favorite Chomsky points. I don't remember the exact phrasing, but something like: "Language is innate in humans because in every household practically all children learn it while none of the pets do."
A similar insight:
We teach children to read and write, but we don't need to teach them to listen and speak.
Do children really lack input data though? Human sensory input is quite a lot of data. Our languages may have common structure because that structure reflects causality and physics encountered by directly sampling the real world.
Keep this as a secret! You are over-scientific and over-intellegent against humanities folks gethered here!
That’s not in the hacker news spirit :( especially ironic to shit talk “humanities folks” with a basic misspelling in your comment. We’re all just trying to reach the truth!
While I think some of the points in the article are interesting, the usual evidence for the Chomskian approach is the relative lack of input data for learning language by children in the wild. How much input data is used to train modern language models?
Do children really lack input data though? Human sensory input is quite a lot of data. Our languages may have common structure because that structure reflects causality and physics encountered by directly sampling the real world.
Your hypothesis is plausible, but less likely than Chomsky’s IMO. One basic version of the typical chomskian response to this point is: many animals have the same input data, how come were the only ones who have evolved any capacity for language at all? The very best animals at language are apes, and we have to successfully teach one concepts that humans learn while still wearing diapers
While I think some of the points in the article are interesting, the usual evidence for the Chomskian approach is the relative lack of input data for learning language by children in the wild. How much input data is used to train modern language models?
1. No Language model is yet even close to the scale of the human brain 2. Depending on what exactly you're trying to teach (perfect grammar, paragraphs of coherent text, basic reasoning), much less data is needed. https://arxiv.org/abs/2305.07759 3. Brains don't start at 0. Evolution, dna/rna etc. There's obviously some pre disposition for language learning in humans but that alone isn't enough ground for a "universa…
> Brains don't start at 0. Evolution, dna/rna etc. There's obviously some pre disposition for language learning in humans but that alone isn't enough ground for a "universal grammar"
That is literally what universal grammar is. All that is left is to argue about the size and content of UG.
Wordcels think LLMs imitate the human brain, when a shape rotator knows they really just imitate human language.
doesn't this make LLM a dead end towards AGI and mostly just a neat specific trick?
I think AGI is a questionable concept. We still don't have a good definition of what intelligence really is, and some people keep moving the goal posts. What we need is AI that fills specific needs we have.
Because Chomsky is trying to build a bird, and Norvig is trying to build an airplane. It's much easier to build an airplane to fly than a bird. Chomsky is trying to explain how humans create language. LLM are creating language, but not the way humans do. Nothing about this paper refutes Chomsky's claims.
Chomsky has been adding parameters to his theory to handle exceptions in a way that mimics the endless series of conditional statements appended to knowledge systems of yore. In neuroscience, predictive processing has gained immense favor and can explain language in ways that have nothing to do with innate grammar. https://en.wikipedia.org/wiki/Predictive_coding Exactly how well did "building a bird" work for buildin…
What are these parameters Chomsky has been adding? I'm interested, not to be taken as some defensive remark or the like.
This paper completly misses the point of linguistics as a discipline that generative grammar operates in. We have known that it is possible to understand language without innateness. That is what linguists do. If you look at how linguists know about innate features, the answer is almost always by first discovering them explicitly while analysing language data; not by opening a brain to see what is innately inside. [0…
For a moment I was going to waste my afternoon arguing with people desperately predisposed to being the underdog in the fight against the Father of Linguistics, but you’ve said everything I ever could beautifully, so if this doesn’t help nothing will. Especially love the last paragraph, clarified that pattern for me. On a lighter note, I do expect “Modern Language Models Refute…” to be the new “All you need is…”! It’…
> For a moment I was going to waste my afternoon arguing with people desperately predisposed to being the underdog in the fight against the Big Mean Socialist
This out-of-the-blue accusation sounds like a confession of your true motives in this conversation: You like the man's politics, so you feel compelled to defend him in an unrelated topic.
> 90% of human communication is unspoken. This is as scientific as the idea of humans just using 10% of our brains.
Of course it is; it’s just a comment on a social media forum. There’s just as little science language motivates me to work. Most of the language society relies on is hallucinations; fiat currency, nation states, constructs like “Senate” and Congress, corporatism, brands, copy-paste of historical terminology, not evidence they’re immutable features of reality. What we recite has nothing to do with what we are. I find…
> The fact that you can replicate coherent text from probabilistic analysis and modeling of a very large corpus does not mean that humans acquire and generate language the same way. Also, the LLMs are cheating! They learned from us. It's entirely possible that you do need syntax/semantics/sapience to create the original corpus, but not to duplicate it. Let's see an AlphaZero-style version of an LLM, that learns langu…
A kind of corollary that I'm sure others have thought of: if llms are so smart and human thought is nothing more than a big language model, why can't they (llms) make up their own training data. Any discussion about how they are "thinking" the way we do is BS, I don't know how so many people who know better have been conned.
> I don't know how so many people who know better have been conned.
The simple reason is: because they don't actually "know better". Maybe they are knowledgeable and skilled in some area, but that doesn't mean they are knowledgeable and skilled in everything.