Live data from Hacker News

Simply explained: How does GPT work?

confusedbit.dev

211–220 of 392 posts

Re: Simply explained: How does GPT work?

#211

Earlier quoted context omitted.

> That's highly reductive of our capacities. I'm not saying that GPT4 is as capable as a human-- it can not be, by design, because its architecture lacks memory/feedback paths that we have. What I'm saying is that HOW it thinks might already be quite close in essence to how WE think. > We are not weighted transformers that can be explained in an arxiv paper. GPT, at the end of the day, is a statistical inference mode…

> What I'm saying is that HOW it thinks might already be quite close in essence to how WE think. How would one be able to prove this? Nobody knows how we think, yet. All one can say is that what GPT-4 outputs could plausible fool another human into believing another human wrote it. But that's exactly what it's designed to do, so what's interesting about that? > If you concede that our brain is "simulatable", It could…

"Nobody knows how we think, yet."

Then how can you confidently say we don't think 'like' Transformers/Attention/Statistical models/etc/etc?

Re: Simply explained: How does GPT work?

#212

On the other hand, many people who are not ready to change, who do not have the skills or who cannot afford to reeducate are threatened. That's me. After programming since the '80s, I'm just so tired. So much work, so much progress, so many dreams lived or shattered. Only to end up here at this strange local maximum, with so much potential, destined to forever run in place by the powers that be. The fundamentals form…

It sounds like your mindset is the root of your struggles. Embracing change and adapting to new technologies has always been crucial in our industry. Instead of waiting for help from others, take control and collaborate with like-minded people. If you don't like the status quo, work toward changing it.

Re: Simply explained: How does GPT work?

#213

I asked it which was better, Lisp or Almonds. It said that was an impossible comparison like Apples and Oranges. Then I asked it which were more similar, Apples & Oranges or Lisp & Almonds. It said it is impossible to classify either of those two pairs as more similar because they too fundamentally different. It couldn't come up with anything like Lisp is not edible. Or that Apples and Oranges are both sweet and Lisp…

I hear people with these outrageous examples and I have to wonder how or what you're actually trying? I just asked GPT-4 (via chat.openai.com):

> Which is more similar, Apples & Oranges or Lisp & Almonds?

and it said:

> When comparing these two pairs, the more similar pair would be Apples & Oranges. Both apples and oranges are fruits, and they share many characteristics, such as being edible, containing vitamins and nutrients, and growing on trees.

> On the other hand, Lisp & Almonds are quite dissimilar. Lisp is a programming language used for computing and software development, while almonds are a type of nut commonly consumed as a food item. They belong to completely different categories and share almost no common characteristics.

And the question about which is better had a reasonable answer, too. This is GPT-3.5 but GPT-4 was roughly equivalent:

> It's not really possible to compare the relative merits of Lisp and almonds, as they are fundamentally different things.

> Lisp is a programming language, and like any programming language, it has its own strengths and weaknesses, depending on what you're trying to accomplish. Some programmers might prefer Lisp for certain tasks, while others might prefer a different language.

> Almonds, on the other hand, are a type of nut that are known for their nutritional value and health benefits. They are a good source of healthy fats, protein, and fiber, and can be a valuable part of a healthy diet.

> So, whether Lisp or almonds are "better" really depends on what you're looking for. If you're a programmer, Lisp might be better for certain tasks, while if you're looking for a nutritious snack, almonds might be a better choice.

Re: Simply explained: How does GPT work?

#214

On the other hand, many people who are not ready to change, who do not have the skills or who cannot afford to reeducate are threatened. That's me. After programming since the '80s, I'm just so tired. So much work, so much progress, so many dreams lived or shattered. Only to end up here at this strange local maximum, with so much potential, destined to forever run in place by the powers that be. The fundamentals form…

It was the best of times, it was the worst of times...

In the long run tech does a bit too well with "food in their belly" to the point that obesity is the main problem in the English speaking world.

As to programming it's quite cool getting chat GTP to write code and stuff. If you can't beat it make use of it I guess.

Re: Simply explained: How does GPT work?

#215
post #166

Earlier quoted context omitted.

> most of those sentences are meaningless so they won't come up in normal use Feel free to come up with a better entropy model then. Stackoverflow gives me confidence that it will be between 5 and 11 bits per word anyway [ https://linguistics.stackexchange.com/questions/8480/what-is... ]. > if statements can grab patterns just fine in most languages, they're not limited to pure equality This does not help you one bit…

> What is the entropy per word of random yet grammatical text? More colourless green dreams sleep furiously in garden path sentences than I have > This does not help you one bit. Dunno, how many bits does ELIZA? I assume more than 1…

> What is the entropy per word of random yet grammatical text?

That is what these 5-11bit estimates are about. Those would correspond to a choice out of 32 to 2048 options (per word), which is much less than there are words in english (active vocabulary for a native speaker should be somewhere around 10000-ish).

Just consider the XKCD "thing explainer" which limits itself to a 1k word vocabulary and is very obviously not idiomatic.

If you want your big if to produce credible output, there is simply no way around the entropy bounds in input and desired output, and those bounds render the concept absolutely infeasible even for I/O lengths of just a few sentences.

Eliza is not comparable to GPT because it does not even hold up to very superficial scrutiny; its not really capable of even pretending to intelligently exchange information with the user, it just relies on some psychological tricks to somewhat keep a "conversation" going...

Re: Simply explained: How does GPT work?

#216
post #135

I’d be interested in hearing from anyone who takes the Chinese Room scenario seriously, or at least can see how it applies to any of this. I cannot see that it matters if a computer understands something. If it quacks like a duck and walks like a duck, and your only need is for it to quack and walk like a duck, then it doesn’t matter if it’s actually a duck or not for all intents and purposes. It only matters if you…

In my understanding of the Chinese Room example, the resolution to the argument is that the *human* may not understand Chinese, but the *system as a whole* can be said to understand it. With this in mind, I think asking whether ChatGPT *in and of itself* is "conscious" or has "agency" is sort of like asking if the speech center of a particular human's brain is "conscious" or has "agency": it's not really a question t…

[deleted]

Re: Simply explained: How does GPT work?

#217

I’d be interested in hearing from anyone who takes the Chinese Room scenario seriously, or at least can see how it applies to any of this. I cannot see that it matters if a computer understands something. If it quacks like a duck and walks like a duck, and your only need is for it to quack and walk like a duck, then it doesn’t matter if it’s actually a duck or not for all intents and purposes. It only matters if you…

Here's GPT output that demonstrates that it doesn't reflect on its text: «GPT Chat: You are right, I apologize for the confusion in my previous responses. The corrected explanation is:

The pattern is counting the number of closed spaces in each letter of the spelled-out number. A closed space is any enclosed space in a letter, such as in the letters "a", "b", "d", "e", etc.

Following the pattern:

- one -> 2 (there are closed spaces in the letters "n" and "e") - two -> 1 (there is a closed space in the letter "o") - three -> 2 (there are closed spaces in the letters "h" and "e") - four -> 1 (there is a closed space in the letter "o") - five -> 1 (there is a closed space in the letter "e") - six -> 0 (there are no closed spaces in the letters) - seven -> 2 (there are closed spaces in the letters "e" and "n") - eight -> 1 (there is a closed space in the letter "g") - nine -> 1 (there is a closed space in the letter "e") - ten -> 1 (there is a closed space in the letter "b") - eleven -> 3 (there are closed spaces in the letters "e", "l", and "v") - twelve -> 2 (there are closed spaces in the letters "b" and "d") - thirteen -> 2 (there are closed spaces in the letters "b" and "d")

Each item follows the pattern, as the number of closed spaces in their letters matches the corresponding number in the pattern.

The whole sequence is:

one -> 2 two -> 1 three -> 2 four -> 1 five -> 1 six -> 0 seven -> 2 eight -> 1 nine -> 1 ten -> 1 eleven -> 3 twelve -> 2 thirteen -> 2 ...»

It is clear the model doesn't know what it is talking about.

Re: Simply explained: How does GPT work?

#218
post #155

Earlier quoted context omitted.

He's got a habit of self aggrandizing, antagonism, and deception in an effort to promote himself and his brand, I worry that his explanations are designed to maximally benefit him, rather than to maximally explain the topic. He's a brilliant man, I just don't trust him.

Is that the case with this specific article?

That’s not how trust works.

Re: Simply explained: How does GPT work?

#219
post #135

I’d be interested in hearing from anyone who takes the Chinese Room scenario seriously, or at least can see how it applies to any of this. I cannot see that it matters if a computer understands something. If it quacks like a duck and walks like a duck, and your only need is for it to quack and walk like a duck, then it doesn’t matter if it’s actually a duck or not for all intents and purposes. It only matters if you…

In my understanding of the Chinese Room example, the resolution to the argument is that the *human* may not understand Chinese, but the *system as a whole* can be said to understand it. With this in mind, I think asking whether ChatGPT *in and of itself* is "conscious" or has "agency" is sort of like asking if the speech center of a particular human's brain is "conscious" or has "agency": it's not really a question t…

Good point, that very much vibes with my thoughts on this matter. Lately, I've been contemplating the analogy between the role LLMs might take within society with that of the brain's language center* in human behavior. There's definitely a way in which we resemble these models. More than some might like to admit. The cleverness, but also the hallucinating, gaslighting and other such behaviors.

And on the other hand, any way you'd slice it, it seems to me LLMs - and software systems in general - necessarily lack intrinsic motivation. By definition, any goal it has can only be the goal of whoever designed that system. Even if its maker decides - "let it pick goals randomly", those randomly picked goals are just intermediate steps toward the enacting of the programmer's original goal. Robert Miles' YouTube videos on alignment shed light on these issues also. For example: https://www.youtube.com/watch?v=hEUO6pjwFOo

Another relevant source on these issues is the book "The Master and his Emissary", which discusses how basically the language center can, in some way - I'm simplifying a lot, fall prey to the illusion that "it" is the entirety of human consciousness.

* or at least some subsystems of that language center, it's important to remember how little we still understand of human cognition

Re: Simply explained: How does GPT work?

#220
post #200

I’d be interested in hearing from anyone who takes the Chinese Room scenario seriously, or at least can see how it applies to any of this. I cannot see that it matters if a computer understands something. If it quacks like a duck and walks like a duck, and your only need is for it to quack and walk like a duck, then it doesn’t matter if it’s actually a duck or not for all intents and purposes. It only matters if you…

Here's an example that I think garners more agreement that properties of a limit ("really understanding") don't necessarily mean that any path towards that limit has the properties of the limit. I think there's a lot of room for disagreement about whether this is a factually-accurate analogy and I'm not trying to argue either way on that, just trying to answer your question about how one might make these sorts of arg…

I think what this points towards is that we care about the internal mechanism. If we prod it externally and it gives the wrong answer, then the internal mechanism is definitely wrong. But if we get the right answers and then open it up and find the internals are still wrong, it's still wrong.

This illuminates a contradiction: the walks like a duck thing is incompatible with the internals being a duck. If you see a creature with feathers that waddles and can fly, it might still be a robot when you open it. So your test cannot just rely on external tests. But you also want to create a definition of artificial intelligence that doesn't depend on being made of meat and electricity.

Post reply on HN