Live data from Hacker News

Chomsky on what ChatGPT is good for (2023)

chomsky.info

141–150 of 389 posts

Re: Chomsky on what ChatGPT is good for (2023)

#141

I confess my opinion of Noam Chomsky dropped a lot from reading this interview. The way he set up a "Tom Jones" strawman and kept dismissing positions using language like "we'd laugh", "total absurdity", etc. was really disappointing. I always assumed that academics were only like that on reddit, and in real life they actually made a serious effort at rigorous argument, avoiding logical fallacies and the like. Yet he…

"Tom Jones" isn't a strawman, Chomsky is addressing an actual argument in a published paper from Steven Piantadosi. He's using a pseudonym to be polite and not call him out by name. > instead of even attempting to summarize the arguments for his position.. He makes a very clear, simple argument, accessible to any layperson who can read. If you are studying insects what you are interested in is how insects do it not w…

>The systems work just as well with impossible languages that infants cannot acquire as with those they acquire quickly and virtually reflexively.

Where is the research on impossible language that infants can't acquire? A good popsci article would give me leads here.

Even assuming Chomsky's claim is true, all it shows is that LLMs aren't an exact match for human language learning. But even an inexact model can still be a useful research tool.

>That’s highly unlikely for reasons long understood, but it’s not relevant to our concerns here, so we can put it aside. Plainly there is a biological endowment for the human faculty of language. The merest truism.

Again, a good popsci article would actually support these claims instead of simply asserting them and implying that anyone who disagrees is a simpleton.

I agree with Chomsky that the postmodern critique of science sucks, and I agree that AI is a threat to the human race.

Re: Chomsky on what ChatGPT is good for (2023)

#142
post #99

The fact that we have figured out how to translate language into something a computer can "understand" should thrill linguists. Taking a word (token) and abstracting it's "meaning" as a 1,000-dimension vector seems like something that should revolutionize the field of linguistics. A whole new tool for analyzing and understanding the underlying patterns of all language! And there's a fact here that's very hard to disp…

Restricted to linguistics, LLM's supposed lack of understanding should be a non-sequitur. If the question is whether LLMs have formed a coherent ability to parse human languages, the answer is obviously yes. In fact not just human languages, as seen with multimodality the same transformer architecture seems to work well to model and generate anything with inherent structure.

I'm surprised that he doesn't mention "universal grammar" once in that essay. Maybe it so happens that humans do have some innate "universal grammar" wired in by instinct but it's clearly not _necessary_ to be able to parse things. You don't need to set up some explicit language rules or generative structure, enough data and the model learns to produce it. I wonder if anyone has gone back and tried to see if you can extract out some explicit generative rules from the learned representation though.

Since the "universal grammar" hypothesis isn't really falsifiable, at best you can hope for some generalized equivalent that's isomorphic to the platonic representation hypothesis and claim that all human language is aligned in some given latent representation, and that our brains have been optimized to be able to work in this subspace. That's at least a testable assumption, by trying to reverse engineer the geometry of the space LLMs have learned.

Re: Chomsky on what ChatGPT is good for (2023)

#143

Earlier quoted context omitted.

> After reading it seems even more true that Chomsky's Tom Jones is a strawman. Lol. It's clear you are not interested in having any kind of rational discussion on the topic and are driven by some kind of zealotry when you claim to have read a technical 40 page paper (with an additional 18 pages of citations) in 30 minutes. Even if by some miraculous feat you had read it you haven't made a single actual argument or a…

It’s certainly not a dense paper with careful nuanced derivations that you have to ponder to grasp. It’s a light read you can skim especially if you aren’t interested in LLM Trump improv and you are familiar with the general thought behind connectionism, construction grammar, other modern linguistic theories and, of course, universal grammar. The debate is as old as UG, but now with a new LLM flavor. I don’t know whi…

> I read it and found nothing similar to “Stop wasting your time; naval vessels do it all the time.”

You implied in the previous paragraph that you didn't in fact read it and you only "skimmed" it. Maybe that's why you "found nothing similar to 'stop wasting your time; naval vessels do it all the time". But even in skimming the paper it's incomprehensible how you could miss it: At least the first 23 pages of the draft version I have just describe how well LLMs perform and completely ignores the relevant question of how human language works. (It doesn't get any better after the first 23 pages). So presumably you just don't know what an analogy is and are literally searching for the term "naval vessels".

Here's just one example demonstrating that Piantodosi does in fact claim what Chomsky says he does: Piantodosi writes "The success of large language models is a failure for generative theories because it goes against virtually all of the principles these theories have espoused." Rewriting that statement using Chomsky's analogy illustrates how idiotic the original statement is: "The success of naval vessels is a failure for insect navigation theories because it goes against all of the principles these theories have espoused".

Re: Chomsky on what ChatGPT is good for (2023)

#144

Earlier quoted context omitted.

"Tom Jones" isn't a strawman, Chomsky is addressing an actual argument in a published paper from Steven Piantadosi. He's using a pseudonym to be polite and not call him out by name. > instead of even attempting to summarize the arguments for his position.. He makes a very clear, simple argument, accessible to any layperson who can read. If you are studying insects what you are interested in is how insects do it not w…

>The systems work just as well with impossible languages that infants cannot acquire as with those they acquire quickly and virtually reflexively. Where is the research on impossible language that infants can't acquire? A good popsci article would give me leads here. Even assuming Chomsky's claim is true, all it shows is that LLMs aren't an exact match for human language learning. But even an inexact model can still…

> Where is the research on impossible language that infants can't acquire? A good popsci article would give me leads here.

It's not infants, it's adults but Moro "Secrets of Words" is a book that describes the experiments and is aimed at lay people.

> Even assuming Chomsky's claim is true, all it shows is that LLMs aren't an exact match for human language learning. But even an inexact model can still be a useful research tool.

If it is it needs to be shown, not assumed. Just as you wouldn't by default assume that GPS navigation tells you about insect navigation (though it might somehow).

> Again, a good popsci article would actually support these claims instead of simply asserting them and implying that anyone who disagrees is a simpleton.

He justifies the statement in the previous sentence (which you don't quote) where he says that it is self-evident by virtue of the fact that something exists at the beginning (i.e. it's not empty space). That's the "merest truism". No popsci article is going to help understand that if you don't already.

Re: Chomsky on what ChatGPT is good for (2023)

#146

Earlier quoted context omitted.

This is where I'm stuck. For other commentators, as I understand it, Chomsky's talking about well-defined grammar and language and production systems. Think Hofstadter's Godel Escher Bach. Not "folk" understanding of language. I have no understanding or intuition, or even a finger nail grasp, for how an LLM generates, seemingly emulating, "sentences", as though created with a generative grammar. Is any one comparing…

I don't really understand your question but if a deep neural network predicts the weather we don't have any problem accepting that the deep neural network is not an explanatory model of the weather (the weather is not a neural net). The same is true of predicting language tokens.

> is not an explanatory model of the weather (the weather is not a neural net)

I don't follow. Aren't those entirely separate things? The most accurate models of anything necessarily account for the underlying mechanisms. Perhaps I don't understand what you mean by "explanatory"?

Specifically in the case of deep neural networks, we would generally suppose that it had learned to model the underlying reality. In effect it is learning the rules of a sufficiently accurate simulation.

Re: Chomsky on what ChatGPT is good for (2023)

#147

Earlier quoted context omitted.

It’s certainly not a dense paper with careful nuanced derivations that you have to ponder to grasp. It’s a light read you can skim especially if you aren’t interested in LLM Trump improv and you are familiar with the general thought behind connectionism, construction grammar, other modern linguistic theories and, of course, universal grammar. The debate is as old as UG, but now with a new LLM flavor. I don’t know whi…

> I read it and found nothing similar to “Stop wasting your time; naval vessels do it all the time.” You implied in the previous paragraph that you didn't in fact read it and you only "skimmed" it. Maybe that's why you "found nothing similar to 'stop wasting your time; naval vessels do it all the time". But even in skimming the paper it's incomprehensible how you could miss it: At least the first 23 pages of the draf…

There is a difference between supporting one research paradigm over another and rejecting science altogether to focus on engineering. The first quote and the context around it implies the latter.

The success of naval vessels shows it’s possible to navigate without innate star and wind comprehension, so maybe we should think of that inner stuff as phlogiston. (Yeah, this analogy isn’t as nice but it’s quite hard to translate the nuance of linguistic debate into nautical terms.)

Re: Chomsky on what ChatGPT is good for (2023)

#148

Earlier quoted context omitted.

I don't really understand your question but if a deep neural network predicts the weather we don't have any problem accepting that the deep neural network is not an explanatory model of the weather (the weather is not a neural net). The same is true of predicting language tokens.

> is not an explanatory model of the weather (the weather is not a neural net) I don't follow. Aren't those entirely separate things? The most accurate models of anything necessarily account for the underlying mechanisms. Perhaps I don't understand what you mean by "explanatory"? Specifically in the case of deep neural networks, we would generally suppose that it had learned to model the underlying reality. In effect…

> The most accurate models of anything necessarily account for the underlying mechanisms

But they don't necessarily convey understanding to humans. Prediction is not explanation.

There is a difference between Einstein's General Theory of Relativity and a deep neural network that predicts gravity. The latter is virtually useless for understanding gravity (that's even if makes better predictions).

> Specifically in the case of deep neural networks, we would generally suppose that it had learned to model the underlying reality. In effect it is learning the rules of a sufficiently accurate simulation.

No, they just fit surface statistics, not underlying reality. Many physics phenomena were predicted using theories before they were observed, they would not be in the training data even though they were part of the underlying reality.

Re: Chomsky on what ChatGPT is good for (2023)

#150

Earlier quoted context omitted.

> is not an explanatory model of the weather (the weather is not a neural net) I don't follow. Aren't those entirely separate things? The most accurate models of anything necessarily account for the underlying mechanisms. Perhaps I don't understand what you mean by "explanatory"? Specifically in the case of deep neural networks, we would generally suppose that it had learned to model the underlying reality. In effect…

> The most accurate models of anything necessarily account for the underlying mechanisms But they don't necessarily convey understanding to humans. Prediction is not explanation. There is a difference between Einstein's General Theory of Relativity and a deep neural network that predicts gravity. The latter is virtually useless for understanding gravity (that's even if makes better predictions). > Specifically in the…

> No, they just fit surface statistics, not underlying reality.

I would dispute this claim. I would argue that as models become more accurate they necessarily more closely resemble the underlying phenomena which they seek to model. In other words, I would claim that as a model more closely matches those "surface statistics" it necessarily more closely resembles the underlying mechanisms that gave rise to them. I will admit that's just my intuition though - I don't have any means of rigorously proving such a claim.

I have yet to see an example where a more accurate model was conceptually simpler than the simplest known model at some lower level of accuracy. From an information theoretic angle I think it's similar to compression (something that ML also happens to be almost unbelievably good at). Related to this, I've seen it argued somewhere (I don't immediately recall where though) that learning (in both the ML and human sense) amounts to constructing a world model via compression and that rings true to me.

> Many physics phenomena were predicted using theories before they were observed

Sure, but what leads to those theories? They are invariably the result of attempting to more accurately model the things which we can observe. During the process of refining our existing models we predict new things that we've never seen and those predictions are then used to test the validity of the newly proposed models.

Post reply on HN