Live data from Hacker News

Simply explained: How does GPT work?

confusedbit.dev

321–330 of 392 posts

Re: Simply explained: How does GPT work?

#321
post #20

Is it possible that we don’t truly know how it works? That there is some emergent behavior inside these models that we’ve created but not yet properly described? I’ve read a few of these articles but I’m still not completely satisfied.

I hate being the bearish guy during the hype cycle, but I think a lot of that is just anthropomorphizing it. They fed it TBs of human text, it spits out human text, we think it's humanesque. Of course maybe I'm wrong and it's AGI and it will find this comment and torture me for for insulting it's intelligence.

I’m also quite bearish on all the AI hype but part of my skepticism means I don’t trust that the AI experts actually know everything. I believe there’s a possibility that their invention this time is actually more interesting (in some dimensions) than they understand it to be.

Re: Simply explained: How does GPT work?

#322
post #299
post #283

Earlier quoted context omitted.

That does seem really impressive. But don't you think that it's pretty likely that this, or something phrased slightly differently, appeared in the training data?

> But don't you think that it's pretty likely that this, or something phrased slightly differently, appeared in the training data? I don't think so, but I could be wrong. It's definitely not "likely", see the math below. I base that on the fact that people seemed to spend quite a bit of time trying to find the phrase "the confetti has left the cannon" that GPT-4 phrased. It seems Google search has no records of it be…

> I base that on the fact that people seemed to spend quite a bit of time trying to find the phrase "the confetti has left the cannon" that GPT-4 phrased. It seems Google search has no records of it before then?

Could it be that the expression in some form has been used in languages other than English?

Re: Simply explained: How does GPT work?

#323

Earlier quoted context omitted.

I have tried multiple times to use Chatgpt to generate Unreal c++ code. It does not do. It spits out class names for slate objects, that inherit from other slate objects. Chatgpt doesn't understand inheritance. It just guesses what might fit inside a parameter grouping, and never suggests something with the right class type. For my use case, it has never quacked like a duck, so to speak. It never performed , the word…

Yes, in this instance I understand failings of today (though copilot has a much better hit rate, and at the moment it’s a great augmentation to coding if you treat it like an enthusiastic intern). My question is about the future. The argument goes that a machine can never understand Chinese, even if it is capable of interpreting Chinese and responding to or acting on the input perfectly every time. My reply is that,…

It's hard to keep this theoretical. Yes a machine is just a machine.

Defining a machine to be conscious, allows the individual to soak their mind in code and silicon as a receptacle for their spirit.

It creates a pull into a 'second mind'. Anybody who believes this is likely to invest heavily in the maintenance of new technology.

A 'conscious machine', creates an uneasy feeling that we should work to embed our spirit, knowledge, intellect into flipped bits, like expectant mothers. That we should work for the machine, and to the ends of the machine.

And that machine is somehow defined-to-be or a naturally, consciously alive (to a large or small degree). It is said to have a mind worthy of a person's professional output and it can hold the power of a marginally believable conversation.

While all of these described properties are vaguely plausible, it does nothing to help me understand the meaning of a technology, and only benefits those looking to create a fevor around a new tech product.

Describing chatgpt as a stochastic parrot or chinese room grants me a metaphor or analogy for the inner functions of the tech. It also lets me see, or otherwise guesstimate the products abilities clearly, without the belief-as-marketing hype.

I can take the stochastic parrot metaphor, to an article about LLMs and understand in a couple of days what took years of research to create.

Following the belief of computing as real human intelligence and that human intelligence is fundamentally mathematical, requires on some level submission of your mind to a machine that has it's own goals programmed in by someone else.

This centuries-long process of trying to encode and store all human knowledge behind the secure walls of complex coded signs.. and it's advocates for that process, create a subtle and deep twinge of future melancholy or dread or something. The idea that all written/typed meaning will be accessible only by the spiritual power brokers, and not our sons.

No. On some level, machines are just machines, like an abacus or a weaving loom. It can host concepts in the same way that a weaving loom is 'intelligent'. It holds it's shape, abstractions and functions by the laws of physics/metaphysics and according to my human dictates.

You follow the raven into the computers-are-conscious dream at your own risk. Computers are leaning towards controlling people rather than emancipating them. Leaning very hard in that direction. Do we want that? Freedom of mind and meaning is valueable.

Re: Simply explained: How does GPT work?

#324
post #174

A good article and well articulated! I would change the introduction to be more impartial and not anthropomorphize GPT. It is not smart and it is not skilled in any tasks other than that for which it is designed. I have the same reservations about the conclusion. The whole middle of the article is good. But to then compare the richness of our human experience to an algorithm that was plainly explained? And then to sp…

> it is not skilled in any tasks other than that for which it is designed. But it wasn't designed. It's not a computer program, where one can make confident predictions about its limitations based on the source code. It's a very large black box. It was trained on guessing the next word. Does that fact alone prove that it cannot have evolved certain internal structures during the training? Do you claim that an artific…

>But it wasn't designed. It's not a computer program, where one can make confident predictions about its limitations based on the source code.

It definitely is exactly that. It's not any more special than any other program that you can write. I am not totally sure that what you describe could ever exist at all.

What makes this program "magic" compared to any other program exactly? There is no physical difference between it and a "regular" program. Both of them are a bunch of source code that gets compiled into an executable and ran by the underlying OS and hardware. There is nothing physically different between it and other software.

https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-...

Re: Simply explained: How does GPT work?

#325

I’d be interested in hearing from anyone who takes the Chinese Room scenario seriously, or at least can see how it applies to any of this. I cannot see that it matters if a computer understands something. If it quacks like a duck and walks like a duck, and your only need is for it to quack and walk like a duck, then it doesn’t matter if it’s actually a duck or not for all intents and purposes. It only matters if you…

You're talking about intelligence - GPT doesn't quack at all. It clearly is not a duck.

Re: Simply explained: How does GPT work?

#326
post #233

Earlier quoted context omitted.

> One uses a wet, squishy brain. The other uses silicon chips. Well then, that settles the debate!

My point is that's not a debate anyone is having. No one claims that ChatGPT is human! The claim is merely that ChatGPT is engaging in (non-human) forms of reasoning, abstraction, creativity, and so on, with varying levels of ability. There's a separate debate on whether the brain produces human thoughts in a similar way to ChatGPT's non-human thought. The question here is whether brains are essentially biological LL…

[deleted]

Re: Simply explained: How does GPT work?

#327
post #233

Earlier quoted context omitted.

> One uses a wet, squishy brain. The other uses silicon chips. Well then, that settles the debate!

My point is that's not a debate anyone is having. No one claims that ChatGPT is human! The claim is merely that ChatGPT is engaging in (non-human) forms of reasoning, abstraction, creativity, and so on, with varying levels of ability. There's a separate debate on whether the brain produces human thoughts in a similar way to ChatGPT's non-human thought. The question here is whether brains are essentially biological LL…

The word "thought" means something. When you use it to describe ChatGPT, you have in fact argued "there's no fundamental difference between humans and LLMs."

Re: Simply explained: How does GPT work?

#328
post #308

Earlier quoted context omitted.

Thing is, they can still solve the problem , even if the problem was not one from its training set. And, more importantly, they solve the problem much better if you tell them to reason about it in writing first before giving the final answer.

Yes I know, as I said they are very knowledgeable and in some ways very intelligent. We just need to bear in mind their processing architecture is radically different from our. This makes our intuitions about their abilities highly error prone.

Absolutely. The shoggoth metaphor is extremely apt here.

What I was specifically responding to is the claim that they can only solve certain kinds of problems because those kinds of problems (and their solutions) were in the training set. By now there's plenty of counter-examples of unique problems that are nevertheless solved. At which point I think we do have to call it "understanding" and "reasoning", even as we acknowledge that it is a very alien form of understanding and reasoning that we just barely managed to squeeze into something that kinda sorta feels humanish.

Re: Simply explained: How does GPT work?

#329

Earlier quoted context omitted.

Multi-head attention just means that you're looking at all the words at once rather than only looking at one word at a time, and using that to generate the next word. So instead of using attention only on the last word you also have attention on the penultimate word and the one before that and the one before that, etc. I think it is fairly obvious why this gives better results than say an RNN – you are utilizing cont…

That's not what multi-head attention means. Multi-head attention is the use of learned projection operators to perform attention operations within multiple lower-dimensional subspaces of the network's embedding space, rather than a single attention operation in the full embedding space. E.g. projecting 10 512-D vectors into 80 64-D vectors, attending separately to the 8 sets of 10 embedding projections, then concaten…

How is that different from what I said?

Re: Simply explained: How does GPT work?

#330
post #203

Earlier quoted context omitted.

Almost all people almost never have truly original ideas. When asked to "tell me an idea [you] have never heard before", they will remix stuff they have heard to get something that "feels" like it's new. In some cases they'll actually be wrong and reproduce something they heard and forgot about hearing, but remember the concept. Most of the time, the remix will be fairly superficial. And remixing stuff it has heard b…

Certainly. I mean we've seen all 26 letters before-- ChatGPT is just remixing them. How does one actually measure novelty, without having to know everything first?

The entire strength of large language models like GPT is that they do know a frighteningly good approximation of everything, in terms of having been trained on text written about it.
Post reply on HN