Live data from Hacker News

Simply explained: How does GPT work?

confusedbit.dev

301–310 of 392 posts

Re: Simply explained: How does GPT work?

#301
post #225

Earlier quoted context omitted.

> There’s no need to ascribe any belief that they can evolve, modify themselves, or spontaneously develop intelligence. But neural networks clearly evolve and are modified during training. Otherwise they would never get any better than a random collection of weights and biases, right? Is the claim then that an artificial neural network can never be trained in such a way that it will exhibit intelligent behavior? >> D…

> Is the claim then that an artificial neural network can never be trained in such a way that it will exhibit intelligent behavior? I think it’s not likely a NN can be trained to exhibit any kind of autonomous intelligence. Science has good models and theories of what intelligence is, what constitutes consciousness, and these models are continuing to evolve based on what we find in nature. I don’t doubt that we can t…

> "it’s not a living, breathing, thinking, autonomous organism"

> "autonomous intelligence"

> "what constitutes consciousness"

> "autonomous thought"

In my mind, this is a list of different concepts.

GPT-4 is definitely not living, breathing or autonomous. It doesn't take any actions on its own. It just responds to text.

Can we stay on just the topic of intelligence?

Let's take this narrow definition: "the ability to reason, plan, solve problems, think abstractly, comprehend complex ideas".

> But in none of these cases do I believe these systems are intelligent

It should be possible to measure whether an entity is intelligent just by asking it questions, right?

Let's say we have an unknown entity at the other end of a web interface. We want to decide where it falls on a scale between stochastic parrot and an intelligent being.

What questions about logical reasoning and problem solving can we ask it to decide that?

And where has GPT-4 failed in that regard?

Re: Simply explained: How does GPT work?

#302

Earlier quoted context omitted.

It’s a fallacy to describe what the machine does as “thinking” because that’s only process you know for achieving the same outcome. When you initiate the model with some input where you expect some particular correct output, that means there exists some completed sequence of tokens that is correct—if that weren’t true then you either wouldn’t ask or else you wouldn’t blame the model for being wrong. Now imagine a mac…

Now consider the case when you tell GPT to "think it out loud" before giving you the answer - which, coincidentally, is a well-known trick that tends to significantly improve its ability to produce good results. Is that thinking?

Maybe. Mechanically we might also describe it as causing the model to condition more explicitly on specific tokens derived from the training data rather than the implicit conditioning happening in the raw model parameters. This would tend to more tightly constrain the output space—making a smaller haystack to look for a needle. And leveraging the fact that “next token prediction” implies some consistency with preceding tokens.

It could be thinking, but I don’t think that’s strong evidence that it is thinking.

Re: Simply explained: How does GPT work?

#303
post #297

Earlier quoted context omitted.

It's not clear to me what point you're trying to make. Why do we need an "articulated theory of abstract reasoning" to say that passing the bar exam or writing code for novel, nontrivial tasks requires reasoning? Seems rather obvious.

You are making a claim that there is some attribute of importance. For that claim to be persuasive, it should be supported with an explanation of what that attribute is and is not, and evidence for or against the meeting of those criteria. So far all you have done is say "Look at the text it puts out, isn't that something?" It's just empty excitement, not a well-reasoned argument.

You keep avoiding this question: does passing the bar exam and writing code for novel, nontrivial tasks require reasoning or doesn't it?

You aren't answering because saying no will sound ridiculous. We all know it requires reasoning.

As for an "attribute of importance", I guess that's subjective, but I've used ChatGPT to write code in a few minutes that would have taken me hours of research and implementation. I've shipped that code to thousands of people. That's enough for it to be important to me, even ignoring other applications, but you certainly have the right to remain unimpressed if you so choose.

Re: Simply explained: How does GPT work?

#304

Earlier quoted context omitted.

I read this and can't help but chuckle... To say that we are nowhere being able to have AGI is quite a bold statement. It was after all only a few months ago where many people also believed we were a long way away from ChatGPT-4. The confidence with which you think we are not weighted transformers or statistical inference models is also puzzling. How could you possibly know that? How do you know that that's not preci…

Ah yes, the old: you can’t prove my deity doesn’t exist argument. Puzzling that I don’t share your faith or point of view? Why? The point is to not ascribe properties attributed to a thing we know doesn’t have them. We can teach people how ChatGPT works without getting into pseudo-philosophical babble about what consciousness is and whether humans can be accurately simulated by an LLM with enough parameters.

IMO the big blindside of your argument is that you MUST either accept that some magic happens in human brains (=> which is HARD to reconciliate with a science-inspired world-view), OR that achieving human-level cognitive performance is a pure hardware/software optimization problem.

The thing is that GPT4 already approaches human level cognitive performance in some tasks, which means you need a strong argument for WHY full human-level performance would be out of reach of gradual improvements to the current approach.

On the other hand, a very strong argument could be made that the very first artificial neural networks had the absolutely right ideas and all the improvements over the last ~40 years were just the necessary scaling/tuning for actually approaching human performance levels...

This is also where I have to recommend V Braitenbergs "Vehicles: Experiments in synthetic psychology" (from 1984!) which aged remarkably well and shaped my personal outlook on the human mind more than anything else.

Re: Simply explained: How does GPT work?

#305
post #297

Earlier quoted context omitted.

You are making a claim that there is some attribute of importance. For that claim to be persuasive, it should be supported with an explanation of what that attribute is and is not, and evidence for or against the meeting of those criteria. So far all you have done is say "Look at the text it puts out, isn't that something?" It's just empty excitement, not a well-reasoned argument.

You keep avoiding this question: does passing the bar exam and writing code for novel, nontrivial tasks require reasoning or doesn't it? You aren't answering because saying no will sound ridiculous. We all know it requires reasoning. As for an "attribute of importance", I guess that's subjective, but I've used ChatGPT to write code in a few minutes that would have taken me hours of research and implementation. I've s…

For a human, it takes human reasoning. But a xerox machine can also output the correct answers given the right inputs, which is exactly what you can say about an LLM.

The "attribute of importance" I'm referring to is "rationality". You keep talking about it like it means something but you can't define it beyond "I'm pretty sure this text was made using it".

Does a tape recording of a bird song "know" how to sing like a bird?

Re: Simply explained: How does GPT work?

#306

Earlier quoted context omitted.

Now consider the case when you tell GPT to "think it out loud" before giving you the answer - which, coincidentally, is a well-known trick that tends to significantly improve its ability to produce good results. Is that thinking?

Maybe. Mechanically we might also describe it as causing the model to condition more explicitly on specific tokens derived from the training data rather than the implicit conditioning happening in the raw model parameters. This would tend to more tightly constrain the output space—making a smaller haystack to look for a needle. And leveraging the fact that “next token prediction” implies some consistency with precedi…

I would say that it's very strong evidence that it is thinking, if that "thinking out loud" output affects outputs in ways that are consistent with logical reasoning based on the former. Which is easy to test by editing the outputs before they're submitted back to the model to see how it changes its behavior.

Re: Simply explained: How does GPT work?

#307
post #305

Earlier quoted context omitted.

You keep avoiding this question: does passing the bar exam and writing code for novel, nontrivial tasks require reasoning or doesn't it? You aren't answering because saying no will sound ridiculous. We all know it requires reasoning. As for an "attribute of importance", I guess that's subjective, but I've used ChatGPT to write code in a few minutes that would have taken me hours of research and implementation. I've s…

For a human, it takes human reasoning. But a xerox machine can also output the correct answers given the right inputs, which is exactly what you can say about an LLM. The "attribute of importance" I'm referring to is "rationality". You keep talking about it like it means something but you can't define it beyond "I'm pretty sure this text was made using it". Does a tape recording of a bird song "know" how to sing like…

Those aren't good analogies. An LLM isn't like a xerox machine or a tape recorder. Again, the answers to the bar exam it passed weren't in its training data. Nor was the code it wrote for me.

I'm using the common, colloquial definition of reasoning. I don't think we need an academic treatise to say that passing the bar exam (without copying the answers) or writing code for a novel task requires reasoning.

You're right that we don't fully understand how the LLM is doing this, but that doesn't mean it isn't happening.

Re: Simply explained: How does GPT work?

#308
post #183

Earlier quoted context omitted.

I think it’s undeniable that LLMs encode knowledge, but the way they do so and what their answers imply, compared to what the same answer from a human would imply, are completely different. For example if a human explains the process for solving a mathematical problem, we know that person knows how to solve that problem. That’s not necessarily true of an LLM. They can give such explanations because they have been tra…

Thing is, they can still solve the problem , even if the problem was not one from its training set. And, more importantly, they solve the problem much better if you tell them to reason about it in writing first before giving the final answer.

Yes I know, as I said they are very knowledgeable and in some ways very intelligent. We just need to bear in mind their processing architecture is radically different from our. This makes our intuitions about their abilities highly error prone.

Re: Simply explained: How does GPT work?

#309

Earlier quoted context omitted.

What goals do we have that aren't essentially all boiled down to whatever evolution, genetics, and our environment have sorted of molded into us?

If you subscribe to a purely mechanistic world-view, i.e. computationalism, then yes. But that's a leap of faith I cannot justify taking. It's a matter of faith, because though we cannot exclude the possibility logically, it also doesn't follow necessarily from our experience of life, at least as far as I can see. Yes, so many times throughout the ages, scientists have discovered mechanisms to explain things which we…

> Then, to maintain integrity, one would have to show the same deference to these models as one shows to their fellow human.

Except they’re not even remotely close to anything like human intelligence. As I wrote in another comment they are very capable systems, to the point where in some ways they show some level of elementary understanding, but in many forms of reasoning they are utterly and completely incapable. Assigning human equivalent cognitive status is patently absurd. And yes I am a physicalist and I see no reason why a computer system could not achieve human equivalent cognitive ability. These just aren’t that. They may be an important step towards it though.

Re: Simply explained: How does GPT work?

#310
post #200

Earlier quoted context omitted.

Here's an example that I think garners more agreement that properties of a limit ("really understanding") don't necessarily mean that any path towards that limit has the properties of the limit. I think there's a lot of room for disagreement about whether this is a factually-accurate analogy and I'm not trying to argue either way on that, just trying to answer your question about how one might make these sorts of arg…

The only thing that separates your mechanism for doing addition from what computers actually do is efficiency. Computers can only add numbers up to some fixed size, e.g. 64 bits, and you have to use repetition to add anything larger. Does that mean computers are not "really doing" addition?

There’s a lot more different than efficiency. We can program computers with algorithms capable of computing any possible addition, the limitation being only the memory of the computer and time, not the algorithm itself. Those algorithms are genuinely doing addition in a way that a pre-computed lookup table is not. It’s the difference between computing an addition in your head and just remembering that 2 + 2 is 4.
Post reply on HN