I am highly skeptical of LLMs as a mechanism to achieve AGI, but I also find this paper fairly unconvincing, bordering on tautological. I feel similarly about this as to what I've read of Chalmers - I agree with pretty much all of the conclusions, but I don't feel like the text would convince me of those conclusions if I disagreed; it's more like it's showing me ways of explaining or illustrating what I already belie…
Isn't any formal "proof" or "reasoning" that shows that something cannot be AGI inherently flawed, because we have a hard time formally describing what AGI is anyway. Like your argument: embodiment is missing in LLMs, but is it needed for AGI? Nobody knows. I feel we first have to do a better job defining the basics of intelligence, we can then define what it means to be an AGI, and only then can we prove that someth…
Large models of what? Mistaking engineering achievements for linguistic agency
71–80 of 162 posts
Re: Large models of what? Mistaking engineering achievements for linguistic agency
#72Earlier quoted context omitted.
> LLMs do not have corporeal experience. But it's not obvious that this means that they cannot, a priori, have an "internal" concept of reality, or that it's impossible to gain such an understanding from text. I would argue it is (obviously) impossible the way the current implementation of models work. How could a system which produces a single next word based upon a likelihood and and a parameter called a "temperatu…
Transformer models have been shown to spontaneously form internal, predictive models of their input spaces. This is one of the most pervasive misunderstandings about LLMs (and other transformers) around. It is of course also true that the quality of these internal models depends a lot on the kind of task it is trained on. A GPT must be able to reproduce a huge swathe of human output, so the internal models it picks o…
Please do :)
Re: Large models of what? Mistaking engineering achievements for linguistic agency
#73I am highly skeptical of LLMs as a mechanism to achieve AGI, but I also find this paper fairly unconvincing, bordering on tautological. I feel similarly about this as to what I've read of Chalmers - I agree with pretty much all of the conclusions, but I don't feel like the text would convince me of those conclusions if I disagreed; it's more like it's showing me ways of explaining or illustrating what I already belie…
Re: Large models of what? Mistaking engineering achievements for linguistic agency
#74Re: Large models of what? Mistaking engineering achievements for linguistic agency
#75How would the authors consider a paralyzed individual who can only move their eyes since birth? That person can learn the same concepts as other humans and communicate as richly (using only their eyes) as other humans. Clearly, the paper is viewing the problem very narrowly.
> ...a paralyzed individual who can only move their eyes since birth... I don't think such an individual is possible.
Re: Large models of what? Mistaking engineering achievements for linguistic agency
#76I'm more or less a layperson when it comes to LLMs and this nascent concept of AI, but there's one argument that I keep seeing that I feel like I understand, even without a thorough fluency with the underlying technology. I know that neural nets, and the mechanisms LLMs employ to train and form relational connections, can plausibly be compared to how synapses form signal paths between neurons. I can see how that make…
Re: Large models of what? Mistaking engineering achievements for linguistic agency
#77I am highly skeptical of LLMs as a mechanism to achieve AGI, but I also find this paper fairly unconvincing, bordering on tautological. I feel similarly about this as to what I've read of Chalmers - I agree with pretty much all of the conclusions, but I don't feel like the text would convince me of those conclusions if I disagreed; it's more like it's showing me ways of explaining or illustrating what I already belie…
I think the deeper point of the paper is that you simply cannot generate an intelligent entity by just looking at recorded language. You can create a dictionary, and a map - but one must not mistake this map for the territory.
Re: Large models of what? Mistaking engineering achievements for linguistic agency
#78I am highly skeptical of LLMs as a mechanism to achieve AGI, but I also find this paper fairly unconvincing, bordering on tautological. I feel similarly about this as to what I've read of Chalmers - I agree with pretty much all of the conclusions, but I don't feel like the text would convince me of those conclusions if I disagreed; it's more like it's showing me ways of explaining or illustrating what I already belie…
I like how Chomsky deals with it who doesn't have any spirituality at all, the big degenerate materialist:
As far as I can see all of this [he's speaking about the Loebner Prize and the Turing test in general] is entirely pointless. It's like asking how we can determine empirically whether an aeroplane can fly the answer being if it can fool someone into thinking that it's an eagle under some conditions.
https://youtu.be/0hzCOsQJ8Sc?si=MUXpmIwAzcla9lvK&t=2052
(My transcript)
He's right, you know. It should be possible to tell whether something is intelligent just as easily as it is to say that something is flying. If there are endless arguments about it, then it's probably not intelligent (yet). Conversely, if everyone can agree it is intelligent then it probably is.
Re: Large models of what? Mistaking engineering achievements for linguistic agency
#79I am highly skeptical of LLMs as a mechanism to achieve AGI, but I also find this paper fairly unconvincing, bordering on tautological. I feel similarly about this as to what I've read of Chalmers - I agree with pretty much all of the conclusions, but I don't feel like the text would convince me of those conclusions if I disagreed; it's more like it's showing me ways of explaining or illustrating what I already belie…
>> Again, that conclusion feels wrong to me... but if I'm being honest with myself, I can't point to why, other than to point at some form of dualism or spirituality as the escape hatch. I like how Chomsky deals with it who doesn't have any spirituality at all, the big degenerate materialist: As far as I can see all of this [he's speaking about the Loebner Prize and the Turing test in general] is entirely pointless.…
Because it's not easy to tell whether something is flying. Definitions like that fall apart every time we encounter something out of the ordinary. If you take the criterion of "there's no discussion about it", then you're limiting the definition to that which is familiar, not that which is interesting.
Is an ekranoplan flying? Is an orbiting spaceship flying? Is a hovercraft flying? Is a chicken flapping its wings over a fence flying?
Your criterion would suggest the answer of "no" to any of those cases, even though those cover much of the same use cases as flying, and possibly some new, more interesting ones.
And I don't think an AGI must be limited to the familiar notion of intelligence to be considered an AGI, or, at the very least, to open up avenues that were closed before.
Re: Large models of what? Mistaking engineering achievements for linguistic agency
#80Earlier quoted context omitted.
The crux of the video game analogy seems to be that when you go close to an object, the resolution starts blurring and the illusion gets broken, and there is a similar thing that happens with LLMs (as of today) as well. This is, so far, reasonable based on daily experience with these models. The extension of that argument being made in the paper is that a model trained on language tokens spewed by humans is incapable…
Why are LLMs incapable of reaching that limit? It's very easy to imagine video games getting to that point. We have all the data to see objects right down to the atomic level, which is plenty more than you'd need for a game. It's mostly a matter of compute. Why then should LLMs breakdown if they can at least mimic the smartest humans? We don't need "resolution" beyond that.
Language Identification in the Limit in short tells us that if there is an automaton equivalent to human language then, if it's at most a regular automaton it can be identified ("learned") by a number of positive only examples approaching infinity, and if it's above regular then a number of negative examples approaching infinity is also needed to identify it. Chomsky based his "Poverty of the Stimulus" argument about linguistic nativism (the built-in "language faculty" of humans) on this result, known as Gold's Result after Mark E. Gold who proved it in the setting of Inductive Inference in 1964. Gold's result is not controversial, but Chomsky's use of it has seen no end of criticism, many from the computational linguistics community (including people in it that have been great teachers to me, without having ever met me, like Charniak, Manning and Schutze, and Jurafsky and Martin) [1].
Those critics generally argue that human language can be learned like everything and anything else: with enough data drawn from a distribution assumed identical to the true distribution of the data in the concept to be learned, and allowing a finite amount of error with a given probability, i.e. under Probably Approximately Correct Learning assumptions, the learning setting introduced by Leslie Valiant in 1984, that replaced Inductive Inference and that serves as the theoretical basis of modern statistical machine learning, in the rare cases where someone goes looking for one. Around the same time that Valiant was describing PAC-Learning, Vapnik and Chervonenkis were developing their statistical learning theory behind the Iron Curtain and if you take a machine learning course in school you'll learn about the VC Dimension and wonder what's that got to do with AI and LLMs.
The big question is how relevant is all this to a) human language and b) learning human language with an LLM. Is there an automaton that is equivalent to human language? Is human language PAC-learnable (i.e. from a polynomial number of examples)? There must be some literature on this in the linguistics community, possibly the cognitive science community. I don't see these questions asked or answered in machine learning.
Rather, in machine learning people seem to assume that if we throw enough data and compute at a problem it must eventually go away, just like generals of old believed that if they sacrifice enough men in a desperate assault they will eventually take Elevation No. 4975 [2]. That's of course ignoring all the cases in the past where throwing a lot of data and compute at a problem either failed completely -which we usually don't hear anything about because nobody publishes negative results, ever- or gave decidedly mixed results, or hit diminishing returns; as a big example see DeepMind's championing of Deep Reinforcement Learning as an approach to real world autonomous behaviour, based on the success of the approach in virtual environments. To be clear, that hasn't worked out and DeepMind (and everyone else) has so far failed to follow the glory of AlphaGo and kin with a real-world agent.
So in short, yeah, there's a lot to say that we may never have enough data and compute to achieve a good enough approximation of human linguistic ability with a large language model, or something even larger, bigger, stronger, deeper, etc.
__________________
[1] See: https://languagelog.ldc.upenn.edu/myl/ldc/swung-too-far.pdf for a history of the debate.