Earlier quoted context omitted.
Chomsky: Birds fly by flapping their wings in a specific way while changing the angle in order to create lift and propulsion. This paper: Planes fly, but don’t flap their wings, ergo Chomsky is wrong.
Chomsky was saying specific things had to be in the brain because it was impossible to do things otherwise. LLMs shoot this argument down even if they aren't how the brain does it.
Modern language models refute Chomsky’s approach to language
221–230 of 244 posts
Re: Modern language models refute Chomsky’s approach to language
#222Earlier quoted context omitted.
Haha what ? They can make training data just fine. https://arxiv.org/abs/2305.07759
Only after they have been fed human language, which was the point.
Re: Modern language models refute Chomsky’s approach to language
#223Earlier quoted context omitted.
It’s odd to see people doomwaving two general reasoning engines. It’s especially hard to parse a dark sweeping condemnation based on…people are investing in it? It doesn’t have the right to assign names to things? Idk what the argument is. My most charitable interpretation is “it cant reason abour anything unless we already said it” which is obviously false.
The heavy investment is what makes this truth uncomfortable - it does not make this truth true (or false). The point is not so much that we already said it, more that the patterns it encodes and surfaces when prompted are patterns in the written corpus, not of the underlying reality (which it has never experienced). Much like a list of all the addresses in the US (or wherever) will tell you very little about the actu…
You've never experienced the "underlying reality" either.
Re: Modern language models refute Chomsky’s approach to language
#224Earlier quoted context omitted.
> one of which is an average 14 year old, the other an honors student college freshman The point is that they're not those things. Yes, language models can produce solutions to language tests that a 14 year old could also produce solutions for, but a calculator can do the same thing in the dimension of math - that doesn't make a calculator a 14 year old.
Very surprised to see these confident assertions still
Re: Modern language models refute Chomsky’s approach to language
#225Earlier quoted context omitted.
There's nothing qualitatively less "in the world" about a language model than a human. Yes, a human has more senses, and is doubtless exposed to huge categories of training data that a language model doesn't have access to - but it's false to draw a sharp dichotomy between knowing what an iPhone looks like, and knowing how people talk about iPhones. Consider two people - one, a Papau New Guinea tribesperson from a pr…
LLMs don't have any senses, not merely fewer. LLMs don't have any concepts, not merely named ones. A concept is a sensory-motor technique abstracted into a pattern of thought developed by an animal, in a spatio-temporal environment, for a purpose. LLMs are just literally an ensemble of statistical distributions over text symbols. In generating text, they're just sampling from a compressed bank of all text ever digiti…
Text is an extremely limited input stream, but an input stream nonetheless. We know that animal intelligence works well enough with any of a range of sensory streams, and different levels of emphasis on those streams - humans are somehow functional despite a lack of ultrasonic perception and primitive sense of smell.
And your definition of a concept is quite self-serving... I say that as a mathematician familiar with many concepts which don't map at all to sensory motor experiences.
Re: Modern language models refute Chomsky’s approach to language
#226Earlier quoted context omitted.
Um but you have the example of English. Modern English was based on Middle English, which in turn is based on Old English, but greatly influenced by Norman on account of the invasion, as well as by Norse
Keep playing it back through history and you'll find the first language invented by people. There is no equivalent accomplishment for AI.
Re: Modern language models refute Chomsky’s approach to language
#227Earlier quoted context omitted.
Eh, it's certainly true that we're throwing tremendously more hardware and power at the problem with an LLM than with a toddler; that's only relevant to whether LLMs refute Chomsky to whatever degree his argument relied on hardware or power consumption (explicitly or implicitly) and my impression is that it didn't.
Wait so then what is his argument? Because you can always postulate that a large enough computer can simulate every human and therefore can learn stuff too — thus you don’t need a human to learn language, nyeh! Obviously, all that stuff ChatGPT says about feelings and emotions came from humans writing it!
Therefore, if we're able to do the same thing by simply applying more resources, that would undermine his argument in a way that doing the same thing with a vastly larger corpus (whatever the resources we throw at it) doesn't.
I should note that this is based on recollections from 20+ years ago and no serious engagement with the article at hand, so, uh, appropriate salt.
Re: Modern language models refute Chomsky’s approach to language
#228Earlier quoted context omitted.
Your hypothesis is plausible, but less likely than Chomsky’s IMO. One basic version of the typical chomskian response to this point is: many animals have the same input data, how come were the only ones who have evolved any capacity for language at all? The very best animals at language are apes, and we have to successfully teach one concepts that humans learn while still wearing diapers
Why can't a small neural network do what LLMs do? Same reason animals can't learn human language. They don't have the capacity. LLMs started getting interesting "all the sudden" when they hit a certain scale, just like biological brains.
There's lots of animals that have very complex brains, and that engage in very complex behaviors - the fact that none of them have even a hint of linguistic ability seems like a strong indicator that these faculties are a human-specific evolution, and not just... well IDK how to even sum up an anti-UG/GG view. "Kids just sorta figure it out" I guess?
Although who knows! As Chomsky likes to say, this whole field is in a pre-Galilean state due to the impossibility of conducting comparative studies
EDIT: Oh just realized you were the parent comment. Well I'd say the small NN vs. LLM example still doesn't convince me of the likelihood of your statement as I understand it; it goes without saying that lots of animals are much better at intuitive understandings of physics (well, kinetics at least) than humans. You ever seen those snakes that jump from tree to tree? craziest shit you'll ever see
Re: Modern language models refute Chomsky’s approach to language
#229Earlier quoted context omitted.
LLMs don't have any senses, not merely fewer. LLMs don't have any concepts, not merely named ones. A concept is a sensory-motor technique abstracted into a pattern of thought developed by an animal, in a spatio-temporal environment, for a purpose. LLMs are just literally an ensemble of statistical distributions over text symbols. In generating text, they're just sampling from a compressed bank of all text ever digiti…
There's a pile of work on multimodal inputs to LLMs, generally finding that less training data is needed as image (or other) data is added to training. Text is an extremely limited input stream, but an input stream nonetheless. We know that animal intelligence works well enough with any of a range of sensory streams, and different levels of emphasis on those streams - humans are somehow functional despite a lack of u…
Sensory-motor expression of concepts is primitive, yes, they become abstracted --- and yes the semantics of those abstractions can be abstract. I'm not talking semantics, i'm talking genesis.
How does one generate representations whose semantics are the structure of the world? Not via text token frequency, this much is obvious.
I dont think the thinnest sense of "2 + 2 = 4" being true is what a mathematician understands -- they understand, rather, the object 2, the map `+` and so on. That is, the proposition. And when they imagine a sphere of radius 4 containing a square of length 2, etc. -- I think there's a 'sensuous, mechanical, depth' that enables and permeates their thinking.
The intellect is formal only in the sense that, absent content, it has form. That content however is grown by animals at play in their environment.
Re: Modern language models refute Chomsky’s approach to language
#230Earlier quoted context omitted.
Humans learn from the structure of the world -- not the structure of language. LLMs cheat at generating text because they do so via a model of the statistical structure of text. We're in the world , it is us who stipulate the meaning of words and the structure of text. And we stipulate new meanings to novel parts of the world daily . What else is an 'iPhone' etc. ? There's nothing in `i P h o n e` which is at all lik…
> Humans learn from the structure of the world -- not the structure of language. You'd be surprised. Many researchers believe that "knowledge" is inseparable from language, and that language is not associative (labels for the world) but relational . For example, in Relational Frame Theory, human cognition is dependent on bidirectional "frames" that link concepts, and those frames are linguistic in nature. LLMs develo…
It almost doesnt matter what your theory of language is --- any even plausible account will radically depart from the above statistical model. There isn't any theory of language which supposes it's an induction across text tokens.
The problem in this whole discussion is that we know what these statistical models are (models of association in text tokens) -- yet people completely ignore this in favour of saying "it works!".
Well "it works" is NOT an explanatory condition, indeed, it's a terrible one. If you took photographs of the night sky for long enough, you'd predict where all the stars are --- these photos do not employ a theory of gravity to achive these.
LLMs are just photographs of books.
There's a really egregious pseudoscience here that the hype-cycle completely suppresses: we know the statistical form of all ML models. We know that via this mechanism arbitrarily accurate predictions, given arbitrarily relevant data, can be made. We know that nothing in this mechanism is explanatory.
This is trivial. If you video tape everything and play it back you'll predict everything. Photographing things does not impart those photographs the properties of those things -- those serve as a limited assocative model.