Live data from Hacker News

Modern language models refute Chomsky’s approach to language

scholar.google.com

131–140 of 244 posts

Re: Modern language models refute Chomsky’s approach to language

#131
post #71

Earlier quoted context omitted.

The classic example which I think you’re referring to is Nicaraguan Sign Language, which developed organically in Nicaraguan schools for deaf children where neither the children nor the teachers knew any other form of sign language. It’s a fascinating story, a complex fully developed language created by children. Seems to indicate that this is indeed a very inate capability among humans in larger groups: https://en.w…

Yeah because all the other things are innate; visual/spatial awareness, touch, smell, vocalization… Humanoids went millions of years literally learning to navigate 3D space and sense “enough heat, food, water” etc Nomadic tribes had built shared resource depots millennia before language. I can see the color gradients of the trees and feel muscles relax without words. Human language beyond some utilitarian labels just…

> 90% of human communication is unspoken.

This is as scientific as the idea of humans just using 10% of our brains.

Re: Modern language models refute Chomsky’s approach to language

#132

Earlier quoted context omitted.

> The fact that you can replicate coherent text from probabilistic analysis and modeling of a very large corpus does not mean that humans acquire and generate language the same way. Also, the LLMs are cheating! They learned from us. It's entirely possible that you do need syntax/semantics/sapience to create the original corpus, but not to duplicate it. Let's see an AlphaZero-style version of an LLM, that learns langu…

Man I remember when people downplayed AlphaGo because it didn't teach itself unsupervised. "Nothing to see here". Only took them a few months to do AlphaZero.

AlphaZero works on chess, shogi and go and other perfect information games with discreet moves and board s.

LLMs need pairwise linear input that is composed of independent and identically distributed data.

Feed forward neural networks are effectively DAGs thus semi-decidable.

LLM require a corpus, that data is generated by humans and isn't a perfect information game.

If you dig into how many feedforward neural network can be written as a single pairwise linear function in lower dimensions, you can help build an intuition on how they work in higher dimensions that are beyond our ability to visualize.

AlphaZero being able to build a model without access to opening books or endgame tables in perfect information games was an achievement in implementation, it was not a move past existential quantifiers to a universal quantification.

LLMs still need human produced corpus because the search space is much larger than a simple perfect information game. The game board rules were the source of compression for AlphaZero, while human produced text is the source for LLMs.

Neither have a 'common sense' understanding of the underlying data, their results simply fit a finite subset of the data in the same way that parametric regression does.

As there are no accepted definitions for intelligence, mathematics is the only way to understand this.

VC dimensionally and set shattering is probably the most accessible to programming backgrounds if you are interested.

Re: Modern language models refute Chomsky’s approach to language

#133

Earlier quoted context omitted.

> The fact that you can replicate coherent text from probabilistic analysis and modeling of a very large corpus does not mean that humans acquire and generate language the same way. Also, the LLMs are cheating! They learned from us. It's entirely possible that you do need syntax/semantics/sapience to create the original corpus, but not to duplicate it. Let's see an AlphaZero-style version of an LLM, that learns langu…

>Also, the LLMs are cheating! No...they aren't. Humans aren't learning from thin air by any stretch of the imagination.

[deleted]

Re: Modern language models refute Chomsky’s approach to language

#134

Earlier quoted context omitted.

>Also, the LLMs are cheating! No...they aren't. Humans aren't learning from thin air by any stretch of the imagination.

Individual humans aren’t, but we’re talking about the emergent properties of a swarm of humans. 6 gigabytes of dna describes an entity that, if copied 60 times and dropped on an island, will produce a viable language in inly a few decades. We haven’t found a NN architecture with this property for language, but we have for , eg, chess, starcraft, and go

This might be too overarching of a statement but "kind" learning environments like games that have set rules and win conditions are very far from the turing completeness of human language.

Re: Modern language models refute Chomsky’s approach to language

#135
post #80

Earlier quoted context omitted.

Because Chomsky is trying to build a bird, and Norvig is trying to build an airplane. It's much easier to build an airplane to fly than a bird. Chomsky is trying to explain how humans create language. LLM are creating language, but not the way humans do. Nothing about this paper refutes Chomsky's claims.

Chomsky has been adding parameters to his theory to handle exceptions in a way that mimics the endless series of conditional statements appended to knowledge systems of yore. In neuroscience, predictive processing has gained immense favor and can explain language in ways that have nothing to do with innate grammar. https://en.wikipedia.org/wiki/Predictive_coding Exactly how well did "building a bird" work for buildin…

…he has? Isn’t his modern term “minimalism”, where he tries to simplify things as much as possible? Regardless, continuing to study the field in no way implies that he’s backed down or meaningfully evolved his basic theories of Universal Grammars. He’s very much still confident in them.

Re: predictive processing, in what way does that relate to language…? Even if you apply it to language in a way not mentioned in the linked article at all, I don’t see even the rough shape of how it would refute (/be mutually exclusive with) generative grammars. Maybe I’m just missing something because I don’t know much neuro?

Re: building a bird… yeah that’s their point, you don't try to build a bird, you try to study birds. Chomsky cares about what we are, not building machines to do our drudgery. I don’t think I agree entirely with that singular focus, but you see the appeal, no?

Re: Modern language models refute Chomsky’s approach to language

#136

Earlier quoted context omitted.

>Also, the LLMs are cheating! No...they aren't. Humans aren't learning from thin air by any stretch of the imagination.

Humans learn from the structure of the world -- not the structure of language. LLMs cheat at generating text because they do so via a model of the statistical structure of text. We're in the world , it is us who stipulate the meaning of words and the structure of text. And we stipulate new meanings to novel parts of the world daily . What else is an 'iPhone' etc. ? There's nothing in `i P h o n e` which is at all lik…

Structure in the “world”? You mean the stream of “tokens” we ingest?

This just comes down to giving transformers more modalities, not just text tokens.

There is nothing about “2” that conveys any “twoness”, this is true of all symbols.

The token “the text ‘iphone’” and the token “visual/tactile/etc data of iphone observation” are highly correlated. That is what you learn. I don’t know if you call that stipulation, maybe, but an LLM correlates too in its training phase. I don’t see the fundamental difference, only a lot of optimizing and architectural improvements to be made.

Edit: and when I say “a lot”, I mean astronomical amounts of it. Human minds are pretty well tuned to this job, it’ll take some effort to come close.

Re: Modern language models refute Chomsky’s approach to language

#137

Earlier quoted context omitted.

It’s odd to see people doomwaving two general reasoning engines. It’s especially hard to parse a dark sweeping condemnation based on…people are investing in it? It doesn’t have the right to assign names to things? Idk what the argument is. My most charitable interpretation is “it cant reason abour anything unless we already said it” which is obviously false.

> one of which is an average 14 year old, the other an honors student college freshman The point is that they're not those things. Yes, language models can produce solutions to language tests that a 14 year old could also produce solutions for, but a calculator can do the same thing in the dimension of math - that doesn't make a calculator a 14 year old.

Very surprised to see these confident assertions still

Re: Modern language models refute Chomsky’s approach to language

#138
post #94

Earlier quoted context omitted.

> The fact that you can replicate coherent text from probabilistic analysis and modeling of a very large corpus does not mean that humans acquire and generate language the same way. Also, the LLMs are cheating! They learned from us. It's entirely possible that you do need syntax/semantics/sapience to create the original corpus, but not to duplicate it. Let's see an AlphaZero-style version of an LLM, that learns langu…

I just asked an LLM to create a language and provide a demonstration and this is what it said. Call it a stochastic parrot if you want, but I’m pretty sure a linguist can prompt it to properly invent a language. Sure, I can invent a new language for you! Let's call it "Vorin" for the purposes of this demonstration. Vorin is a tonal language with a complex system of noun classes and a relatively simple verb conjugatio…

A small linguistics YouTuber I follow named K Klein made an interesting video (https://www.youtube.com/watch?v=e9NxTi5ZsOo) trying to get ChatGPT to make a language.

The results were... not all that impressive. There were significant issues getting it to consistently apply the rules of the language it had created, even from one prompt to the next -- and after a certain point, it decided to just give Arabic translations instead of the conlang it was supposed to be making up.

Perhaps a more dedicated "prompt engineer"/linguist-type might be able to get better results, but the problem here seems to be similar to the problem trying to get ChatGPT to do arithmetic and other extreme sports. When trying to get it to do anything other than generating one-off syntactically-correct responses to simple prompts in already-existing human languages, it falls down horribly.

Re: Modern language models refute Chomsky’s approach to language

#139
post #48
post #6

While I think some of the points in the article are interesting, the usual evidence for the Chomskian approach is the relative lack of input data for learning language by children in the wild. How much input data is used to train modern language models?

I agree. Question has not been settled. 20 year old human has * heard ~220 million words, talked 50 million words. * read ~10 million words. * experienced 420 million seconds of wakeful interaction with the environment (can be used to estimate the limit to conscious decisions, or number of distinct 'epochs' we experience) From a machine learning perspective human life is surprisingly small set of inputs and actions,…

LLMs learn from unlabeled data. Children definitely do not. There's a huge difference. I would not be surprised if LLMs could learn a lot more efficiently if they had carefully constructed training data, with video and sound.

But also humans have been speaking for so long it's silly to imagine we don't have some evolved language structures in the brain. I don't know why anyone would single that out for skepticism while not questioning e.g. the brain structures for sight, sound, emotions, navigation, etc.

Re: Modern language models refute Chomsky’s approach to language

#140
post #118
post #100

Earlier quoted context omitted.

Also, dogs and cats, and even our close relatives, primates, don't develop a capacity to language.

Yes, and that's one of my favorite Chomsky points. I don't remember the exact phrasing, but something like: "Language is innate in humans because in every household practically all children learn it while none of the pets do."

They[1] do learn some of it, but a fraction so small it doesn't change the argument.

[1]: at least the mamal pets, not goldfishes

Post reply on HN