Live data from Hacker News

Large models of what? Mistaking engineering achievements for linguistic agency

arxiv.org

21–30 of 162 posts

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#21
post #14

I am highly skeptical of LLMs as a mechanism to achieve AGI, but I also find this paper fairly unconvincing, bordering on tautological. I feel similarly about this as to what I've read of Chalmers - I agree with pretty much all of the conclusions, but I don't feel like the text would convince me of those conclusions if I disagreed; it's more like it's showing me ways of explaining or illustrating what I already belie…

>that sufficiently advanced mimicry is not only indistinguishable from the real thing, but at the limit in fact is the real thing. While sufficiently does a lot of the heavy lifting here, the indistinguishable criteria implicitly means there must be no-way to tell if it is not the real thing. The belief that it is the real thing comes from the intuition that anything that can be everything a person must be, but have…

> Future models will not be able to do those things if they are the same as the current ones

I think a lot of people disagree with this. People think if we just keep adding parameters and data, magic will happen. That’s kind of what happened with ChatGPT after all.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#23
post #6

Full paper: [1]. Not much new here. The basic criticism is that LLMs are not embodied; they have no interaction with the real world. The same criticism can be applied to most office work. Useful insight: "We (humans) are always doing more than one thing." This is in the sense of language output having goals for the speaker, not just delivering information. This is related to the problem of LLMs losing the thread of a…

The optimization process adjusts the weights of a computational graph until the numeric outputs align with some baseline statistics of a large data set. There is no "punishment" or "reward", gradient descent isn't even necessary as there are methods for modifying the weights in other ways and the optimization still converges to a desired distribution which people claim is "intelligent".

The converse is that people are "just" statistical distributions of the signals produced by them but I don't know if there are people who claim they are nothing more than statistical distributions.

I think people are confused because they do not really understand how software and computers work. I'd say they should learn some computability theory to gain some clarity but I doubt they'd listen.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#24

Earlier quoted context omitted.

LLMs do contain conceptual representations and LLMs are capable of abstract reasoning. This is trivially provable by asking them to reason about something that is a) purely abstract and b) not in the training data, e.g. "All floots are gronks. Some gronks are klorps. Are any floots klorps?" Any of the leading LLMs will correctly answer questions of this type much more often than chance.

That is not an example of a LLM being capable of abstract reasoning. Changing the question from "What is the capital of United States?" which is easily answerable to something completely abstract and "not in the training model" doesn't change that LLM's are just very advanced text prediction, and always will be. The nature of their design means they are incapable of AGI.

The question I gave is a literal textbook example of abstract reasoning. LLMs are just very advanced text prediction, but they are also provably capable of abstract reasoning. If you think that those statements are contradictory, I would encourage you to read up on the Bayesian hypotheses in cognitive science - it is highly plausible that our brains are also just very advanced prediction models.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#25

I am highly skeptical of LLMs as a mechanism to achieve AGI, but I also find this paper fairly unconvincing, bordering on tautological. I feel similarly about this as to what I've read of Chalmers - I agree with pretty much all of the conclusions, but I don't feel like the text would convince me of those conclusions if I disagreed; it's more like it's showing me ways of explaining or illustrating what I already belie…

> LLMs do not have corporeal experience. But it's not obvious that this means that they cannot, a priori, have an "internal" concept of reality, or that it's impossible to gain such an understanding from text. I would argue it is (obviously) impossible the way the current implementation of models work. How could a system which produces a single next word based upon a likelihood and and a parameter called a "temperatu…

On that point, I would dispute the premise that "it's impossible to have true language skills without implicitly having a representation of self and environment". I don't see any contradiction between the following two ideas:

1. LLMs inherently lack any form of consciousness, subjective experience, emotions, or will

2. A sufficiently advanced LLM with sufficient compute resources would perform on par with human intelligence at any given task, insofar as the task is applicable to LLMs

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#26
I'm more or less a layperson when it comes to LLMs and this nascent concept of AI, but there's one argument that I keep seeing that I feel like I understand, even without a thorough fluency with the underlying technology. I know that neural nets, and the mechanisms LLMs employ to train and form relational connections, can plausibly be compared to how synapses form signal paths between neurons. I can see how that makes intuitive sense.

I'm struggling to articulate my cognitive dissonance here, but is there any empirical evidence that LLMs, or their underlying machine learning technology, share anything at all with biological consciousness beyond a convenient metaphor for describing "neural networks" using terms borrowed from neuroscience? I don't know that it necessarily follows that just because something was inspired by, or is somehow mimicking, the structure of the brain and its basic elements, that it should necessarily relate to its modeled reality in any literal way, let alone provide a sufficient basis for instantiating a phenomena we frankly know very little about. Not for nothing, but our models naturally cannot replicate any biological functions we do not fully understand. We haven't managed to reproduce biological tissues that are exponentially less complex than the brain, are we really claiming that we're just jumping straight past lab-grown t-bones to intelligent minds?

I'm sure most of the people reading this will have seen Matt Parker's videos where they "teach" matchbooks to win a game against humans. Is anyone suggesting those matchbooks, given infinite time and repetition, would eventually spark emergent consciousness?

> The argument would be that that conceptual model is encoded in the intermediate-layer parameters of the model, in a different but analogous way to how it's encoded in the graph and chemical structure of your neurons.

Sorry if I have misinterpreted anyone. I honestly thought all the "neuron" and "synapse" references were handy metaphors to explain otherwise complex computations that resemble this conceptual idea of how our brains work. But it reads a lot like some of the folks in this thread believe it's much more than metaphors, but rather a literal analog.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#27
post #7

That's a lot of thinking they've done about LLMs, but how much did they actually try LLMs? I have long threads where ChatGPT refine solutions to coding problems. Their example of losing the thread after printing a tiny list of 10 philosophers seems really outdated. Also it seems LLMs utilize nested contexts as well, for example when it can break it' own rules while telling a story or speaking hypothetically.

For a paper submitted on July 11, 2024, and with several references to other 2024 publications, it is indeed strange that it gives ChatGPT output from April 2023 to demonstrate that “LLMs lose the thread of a conversation with inhuman ease, as outputs are generated in response to prompts rather than a consistent, shared dialogue” (Figure 1). I have had many consistent, shared dialogues with recent versions of ChatGPT and Claude without any loss of conversation thread even after many back-and-forths.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#28
post #21
post #14

Earlier quoted context omitted.

>that sufficiently advanced mimicry is not only indistinguishable from the real thing, but at the limit in fact is the real thing. While sufficiently does a lot of the heavy lifting here, the indistinguishable criteria implicitly means there must be no-way to tell if it is not the real thing. The belief that it is the real thing comes from the intuition that anything that can be everything a person must be, but have…

> Future models will not be able to do those things if they are the same as the current ones I think a lot of people disagree with this. People think if we just keep adding parameters and data, magic will happen. That’s kind of what happened with ChatGPT after all.

I'm not so sure that view is very widespread amongst people familiar with how LLMs work. Certainly they become more capable with parameters and data, but there are fundamental things that can't be overcome with a basic model and I don't think anyone is seriously arguing otherwise.

For instance LLMs are pretty much stateless without their context window. If you treat the raw generated output as the first and final result then there is very little scope for any advanced consideration of anything.

If you give it a nice long context, give it the ability to edit that context or even access to a key-value function interface, then treat everything it says as internal monologue except for anything in tags which is what the user gets to see. There are plenty of people who see AGI somewhere along that path, but once you take a step down that path, it's no-longer "Just an LLM" the LLM is a component in a greater system.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#29

Earlier quoted context omitted.

I agree that's an argument. I would contend that argument is obviously false. If it were true, LLMs could multiply scalar numbers together trivially. It should be the easiest thing in the world for them. The network required to do that well is extremely small, the parameter sizes of these models are gigantic, and the textual expression is highly regular: multiplication is the simplest concept imaginable. That they ca…

LLMs do contain conceptual representations and LLMs are capable of abstract reasoning. This is trivially provable by asking them to reason about something that is a) purely abstract and b) not in the training data, e.g. "All floots are gronks. Some gronks are klorps. Are any floots klorps?" Any of the leading LLMs will correctly answer questions of this type much more often than chance.

I just asked chatgpt

"All floots are gronks. Some gronks are klorps. Are any floots klorps?"

------

To determine if any floots are klorps, let's analyze the given statements:

1. All floots are gronks. This means every floot falls into the category of gronks. 2. Some gronks are klorps. This means there is an overlap between the set of gronks and the set of klorps.

Since all floots are included in the set of gronks and some gronks are klorps, it is possible that some floots are klorps. However, we cannot conclusively say that any floots are klorps without additional information. It is only certain that if there is any overlap between floots and klorps, it is possible, but not guaranteed, that some floots are klorps.

Re: Large models of what? Mistaking engineering achievements for linguistic agency

#30

Earlier quoted context omitted.

That is not an example of a LLM being capable of abstract reasoning. Changing the question from "What is the capital of United States?" which is easily answerable to something completely abstract and "not in the training model" doesn't change that LLM's are just very advanced text prediction, and always will be. The nature of their design means they are incapable of AGI.

The question I gave is a literal textbook example of abstract reasoning. LLMs are just very advanced text prediction, but they are also provably capable of abstract reasoning. If you think that those statements are contradictory, I would encourage you to read up on the Bayesian hypotheses in cognitive science - it is highly plausible that our brains are also just very advanced prediction models.

You're quite right that LLMs can seemingly do some abstract reasoning problems, but I would not say they aren't in the training data.

Sure, the exact form using the made up word gronk might not be in the training data, but the general form of that reasoning problem definitely exists, quite frequently in fact.

Post reply on HN