Live data from Hacker News

Modern language models refute Chomsky’s approach to language

scholar.google.com

1–10 of 244 posts

Re: Modern language models refute Chomsky’s approach to language

#6
While I think some of the points in the article are interesting, the usual evidence for the Chomskian approach is the relative lack of input data for learning language by children in the wild.

How much input data is used to train modern language models?

Re: Modern language models refute Chomsky’s approach to language

#8
The author (bafflingly) seems to have completely missed the point- since anything they state up to page 15 (at which point I stopped reading) does not refute Chomsky's points at all. The author talks about LLMs and how they generate text and then goes on to talk about how it refutes Chomsky's claim about syntax and semantics. However it does not since Chomsky's primary claim is about how HUMANS acquire language.

The fact that you can replicate coherent text from probabilistic analysis and modeling of a very large corpus does not mean that humans acquire and generate language the same way. [edited page = 15]

Re: Modern language models refute Chomsky’s approach to language

#9
post #6

While I think some of the points in the article are interesting, the usual evidence for the Chomskian approach is the relative lack of input data for learning language by children in the wild. How much input data is used to train modern language models?

Do children really lack input data though? Human sensory input is quite a lot of data. Our languages may have common structure because that structure reflects causality and physics encountered by directly sampling the real world.
Post reply on HN