Live data from Hacker News

Simply explained: How does GPT work?

confusedbit.dev

351–360 of 392 posts

Re: Simply explained: How does GPT work?

#351
post #348

Earlier quoted context omitted.

>But it wasn't designed. It's not a computer program, where one can make confident predictions about its limitations based on the source code. It definitely is exactly that. It's not any more special than any other program that you can write. I am not totally sure that what you describe could ever exist at all. What makes this program "magic" compared to any other program exactly? There is no physical difference betw…

Another example of "evolved behavior" is here, where a robot is trained to walk, run, etc: https://mrl.snu.ac.kr/research/ProjectAgile/Agile.html This is done using neural networks. I believe a project like that can be done by a few researchers over months, not years? If you do this using "regular programming" instead, you'd have to write an insanely complex application that uses inverse kinematics etc. https://en.wi…

There is no such thing as non-regular programming though, that's my point.

All programs that run on the computer have the same "power" in terms of what they can do and what can be computed using them. A program that implements a neural net is not inherently any different than a silly python script. One just does a lot more stuff and is much more interesting.

Re: Simply explained: How does GPT work?

#352
post #348

Earlier quoted context omitted.

Another example of "evolved behavior" is here, where a robot is trained to walk, run, etc: https://mrl.snu.ac.kr/research/ProjectAgile/Agile.html This is done using neural networks. I believe a project like that can be done by a few researchers over months, not years? If you do this using "regular programming" instead, you'd have to write an insanely complex application that uses inverse kinematics etc. https://en.wi…

There is no such thing as non-regular programming though, that's my point. All programs that run on the computer have the same "power" in terms of what they can do and what can be computed using them. A program that implements a neural net is not inherently any different than a silly python script. One just does a lot more stuff and is much more interesting.

Sure, if you drill down, everything is just a Turing machine.

And then we can drill down even further where everything is just physics with atoms, quantum mechanics, etc.

So you're not different from a computer. Both are just physics.

But that's not a useful world view in my opinion.

I think "regular" and "non-regular" programming is a useful distinction.

In regular programming, I have to write explicit implementations of the algorithms in the program.

In "non-regular programming" (neural networks), I just have to know how to set up and train neural networks.

Once I do that, the neural networks can be trained to evolve algorithms that I myself don't know how to implement.

Don't you see the big difference between "I have to code the algorithms" and "the computer does it for me"?

Re: Simply explained: How does GPT work?

#353

Earlier quoted context omitted.

> I struggle to understand why this thing works the way it does. I'm not in this field but have recently found myself going on the deepest dive possible into it as my small brain can absorb. I now know about (on a surface level) neural networks, transformers, attention mechanisms, vectors, maticies, tokenization, loss functions and all sorts of other crazy stuff. I come out of this realizing that there are some incre…

Got a good Youtube list? Other than the HN threads and submissions I can look up this weekend.

this series is good https://www.youtube.com/watch?v=Nw_PJdmydZY

Re: Simply explained: How does GPT work?

#354

Earlier quoted context omitted.

The word "thought" means something. When you use it to describe ChatGPT, you have in fact argued "there's no fundamental difference between humans and LLMs."

The parent was very careful to distinguish "human thought" from "non-human thought".

> The parent was very careful to distinguish "human thought" from "non-human thought".

Yes, I noticed. Putting "non-human" in front of "thought" doesn't help.

I doubt parent uses the word "thought" to describe how a thermostat, calculator, or "Hello world" program works.

Using it to describe ChatGPT has no discernable semantic meaning other than OP believes ChatGPT works like an animal brain.

Re: Simply explained: How does GPT work?

#355

Earlier quoted context omitted.

The word "thought" means something. When you use it to describe ChatGPT, you have in fact argued "there's no fundamental difference between humans and LLMs."

That presupposes that the only thought that exists or can exist is human thought. You can define it that way if you like, but it’s not the only definition.

I'm not saying the only thought that exists is human thought. (I believe animals can think).

I'm saying using a word invented to describe animals behavior, "thought" to describe a large language model has no discernible meaning other than you think it works like an animal brain.

If you think it's an open question whether it works like an animal, you should find a better word than "thought".

Re: Simply explained: How does GPT work?

#356
post #290

Earlier quoted context omitted.

It is a large language model. It manipulates text based on context and the imprint of its vast training. You are not able to articulate a theory of reasoning. You are just pointing to the output of an algorithm and saying "this must mean something!" There isn't even a working model of reasoning here, it's just a human being impressed that a tool for manipulating symbols is able to manipulate symbols after training it…

ttpphd says > "Where is your articulated theory of abstract reasoning?" If he had a complete answer to your questions then he would keep his mouth shut and go directly to META and collect $2 BN USD or get a Nobel prize (or both). What you seem to want is a peer-reviewed academic paper but what we're doing here is brainstorming about what is going on in these LLMs. He's definitely onto something here: LLM models, at t…

I don't like buying into hype mindlessly. I prefer to reason through things and apply skepticism. If people are gonna claim that a chatbot has gained sentience, I'm gonna have some tough questions.

Re: Simply explained: How does GPT work?

#357
post #305

Earlier quoted context omitted.

For a human, it takes human reasoning. But a xerox machine can also output the correct answers given the right inputs, which is exactly what you can say about an LLM. The "attribute of importance" I'm referring to is "rationality". You keep talking about it like it means something but you can't define it beyond "I'm pretty sure this text was made using it". Does a tape recording of a bird song "know" how to sing like…

Those aren't good analogies. An LLM isn't like a xerox machine or a tape recorder. Again, the answers to the bar exam it passed weren't in its training data. Nor was the code it wrote for me. I'm using the common, colloquial definition of reasoning. I don't think we need an academic treatise to say that passing the bar exam (without copying the answers) or writing code for a novel task requires reasoning. You're righ…

Thank you, yes, for saying I am right in saying that the evidence is lacking, which was precisely my original point.

Re: Simply explained: How does GPT work?

#358

I’d be interested in hearing from anyone who takes the Chinese Room scenario seriously, or at least can see how it applies to any of this. I cannot see that it matters if a computer understands something. If it quacks like a duck and walks like a duck, and your only need is for it to quack and walk like a duck, then it doesn’t matter if it’s actually a duck or not for all intents and purposes. It only matters if you…

What I find really entertaining is the "just predicting the next token" argument. If just predicting the next token can produce similar or better results than the almighty human intelligence on some tasks, then maybe there's a bit of hubris in how smart we think we actually are.

That'd be like saying that search engines are smarter than the almighty human intelligence because they know the capitals of every country while most humans don't. No, it just has access to a lot of data near-instantaneously. Just like GPT-4 does. It's the enormity of compiled human knowledge that is "smart" in GPT-4. It absolutely is "just predicting the next token", and it turns out that's enough to be an astoundingly intelligent-seeming system when trained on thousands of years of human knowledge. Of course it is! It's like in Avatar: The Last Airbender when he consults with his thousand past-lives at once for wisdom. GPT-4 lets us consult with the collective knowledge of humanity! It's absolutely amazing! And it's also "just predicting the next token". Those are both true.

Re: Simply explained: How does GPT work?

#359
post #357

Earlier quoted context omitted.

Those aren't good analogies. An LLM isn't like a xerox machine or a tape recorder. Again, the answers to the bar exam it passed weren't in its training data. Nor was the code it wrote for me. I'm using the common, colloquial definition of reasoning. I don't think we need an academic treatise to say that passing the bar exam (without copying the answers) or writing code for a novel task requires reasoning. You're righ…

Thank you, yes, for saying I am right in saying that the evidence is lacking, which was precisely my original point.

The evidence isn’t lacking :) We have lots of evidence. What we lack is a coherent theory that explains the evidence.

Re: Simply explained: How does GPT work?

#360

I’d be interested in hearing from anyone who takes the Chinese Room scenario seriously, or at least can see how it applies to any of this. I cannot see that it matters if a computer understands something. If it quacks like a duck and walks like a duck, and your only need is for it to quack and walk like a duck, then it doesn’t matter if it’s actually a duck or not for all intents and purposes. It only matters if you…

I don't know if I take the Chinese Room scenario seriously, it's a little hard to think about. For me the key is that an LLM (and maybe the Chinese room?) is only trained on text, and therefore its entire "universe" is merely representational. To the LLM there is no difference between the color blue and the word "blue", or a dog and the word "dog". People understand that a word is a provisional marker for something with almost infinite complexity, and when a word fails to meet the goals of the person, the word is discarded, amended, or elaborated until it expresses the thing which has heretofore been unexpressed. For an LLM there is no deeper world to access beyond text.

I don't think an intelligence needs to be human, and it should be physically possible to create an intelligence which is synthetic. But in order to call the intelligence "general", and to rely on it for the purposes that designation implies, it would need to be able to successfully navigate the world, which requires access to that world and the use of the world as its own model, rather than the much simpler and coarser intermediary of text. In order to claim that an LLM can fully navigate the world after being trained on pure text, we would have to believe that all our writings across history have exhausted what there is to say about the world. This is not to say an LLM cannot be useful for some purposes, but there will be key ways in which they fail because they have no sense of meaning or what the world is like. Whether consciousness is required to solve this I don't know, but we simply haven't begun to approach a system that can meaningfully address the world as a world.

Post reply on HN