Live data from Hacker News

Simply explained: How does GPT work?

confusedbit.dev

171–180 of 392 posts

Re: Simply explained: How does GPT work?

#171
post #148

Earlier quoted context omitted.

It's all about emergent complexity. While you can reduce it to "just" statistical auto-completion of the next word, we are seeing evidence of abstraction and reasoning produced as a higher-order effect of these simple completions. It's a bit like the Sagan quote: "If you wish to make an apple pie from scratch, you must first invent the universe". Sometimes for GPT to "just" complete the next word in a way that humans…

> Sometimes for GPT to "just" complete the next word in a way that humans find plausible, it must, along the way, develop a model of the world, theory of mind, abstract reasoning. etc. I did an experiment recently where I asked ChatGPT to "tell me an idea [you] have never heard before". ChatGPT replied with what sounded like an idea for a startup, which was delivering farm-fresh vegetables to customers' doors. This i…

I think it's certainly fair to say that GPT's "reasoning" is different from human reasoning. But I think the core debate we're having is whether the difference really matters in some situations.

Certainly, Midjourney's "creativity" is different from human creativity. But it is producing results that we marvel at. It's creative not because it's doing the exact same philosophical thing humans do, but because it can produce the same effect.

And I think many situations are like that. We can always say that human creativity/reasoning/x will always be different from artificial reasoning. But even today, GPT's statistical model replicates many aspects of human reasoning virtually. Is that really an illusion (implying its fake and potentially useless), or is it just a different way of achieving a similar result?

Plus, different models will excel at different thing. GPT's model will excel at synthesizing answers from far more information than a single human will ever be able to know. Does it really matter if it's not identical to human reasoning on a philosophical or biological level, if it can do things humans can't do?

At the end of the day, some of these discussions feel like bike shedding about what words like "reasoning" mean philosophically. But what will ultimately matter is how well these models perform at real world tasks, and what impact that will have on humanity. It doesn't really matter if it's virtualized reasoning or "real" human reasoning at that point.

Re: Simply explained: How does GPT work?

#172
post #161

Earlier quoted context omitted.

It's all about emergent complexity. While you can reduce it to "just" statistical auto-completion of the next word, we are seeing evidence of abstraction and reasoning produced as a higher-order effect of these simple completions. It's a bit like the Sagan quote: "If you wish to make an apple pie from scratch, you must first invent the universe". Sometimes for GPT to "just" complete the next word in a way that humans…

This is really lofty language without much evidence to back it up. It fluffs up techie people and makes them feel powerful, but it doesn't really describe large language models nor does it describe linguistic processes.

The evidence is ChatGPT's output. Unless you're saying that passing the bar exam, writing working code, etc. doesn't require abstract reasoning abilities or a model of the world?

Re: Simply explained: How does GPT work?

#173

> It is able to link ideas logically, defend them, adapt to the context, roleplay, and (especially the latest GPT-4) avoid contradicting itself. Isn't this just responding to the context provided? Like if I say "Write a Limerick about cats eating rats" isn't it just generating words that will come after that context, and correctly guessing that they'll rhyme in a certain way? It's really cool that it can generate coh…

>Like if I say "Write a Limerick about cats eating rats" isn't it just generating words that will come after that context, and correctly guessing that they'll rhyme in a certain way?

I guess ... this is what confuses me. GPT -- at least, the core functionality of GPT-based products as presented to the end user -- can't just be a language model, can it? There must be vanishingly view examples from its training text that start as "Write a Limerick", followed immediately by some limerick -- most such poems do not appear in that context at all! If it were just "generating some text that's likely to come after that in the training set", you'd probably see some continuations that look more like advice for writing Limericks.

And the training text definitely doesn't have stuff like, "As a language model, I can't provide opinions on religion" that coincides precisely with the things OpenAI doesn't want its current product version to output.

Now, you might say, "okay okay sure, they reach in and tweak it to have special logic for cases like that, but it's mostly Just A Language Model". But I don't quite buy that either -- there must be something outside the language model that is doing significant work in e.g. connecting commands with "text that is following those commands", and that seems like non-trivial work in itself, not reasonably classified as a language model.[2]

If my point isn't clear, here is the analogous point in a different context: often someone will build an AND gate out of pneumatic tubes and say, "look, I made a pneumatic computer, isn't that so trippy? This is what a computer is doing, just with electronics instead! Golly gee, it's so impressive what compressed air is [what LLMs are] capable of!"

Well, no. That thing might count as an ALU[1] (a very limited one), but if you want to get the core, impressive functionality of the things-we-call-computers, you have to include a bunch of other, nontrivial, orthogonal functionality, like a) the ability read and execute a lot of such instructions, and b) to read/write from some persistent state (memory), and c) have that state reliably interact with external systems. Logic gates (d) are just one piece of that!

It seems GPT-based software is likewise solving other major problems, with LLMs just one piece, just like logic gates are just one piece of what a computer is doing.

Now, if we lived in a world where a), b), and c) were well-solved problems to point of triviality, but d) were a frustratingly difficult problem that people tried and failed at for years, then I would feel comfortable saying, "wow, look at the power of logic gates!" because their solution was the one thing holding up functional computers. But I don't think we're in that world with respect to LLMs and "the other core functionality they're implementing".

[1] https://en.wikipedia.org/wiki/Arithmetic_logic_unit?useskin=...

[2] For example, the chaining together of calls to external services for specific types of information.

Re: Simply explained: How does GPT work?

#174

A good article and well articulated! I would change the introduction to be more impartial and not anthropomorphize GPT. It is not smart and it is not skilled in any tasks other than that for which it is designed. I have the same reservations about the conclusion. The whole middle of the article is good. But to then compare the richness of our human experience to an algorithm that was plainly explained? And then to sp…

> it is not skilled in any tasks other than that for which it is designed.

But it wasn't designed. It's not a computer program, where one can make confident predictions about its limitations based on the source code.

It's a very large black box. It was trained on guessing the next word. Does that fact alone prove that it cannot have evolved certain internal structures during the training?

Do you claim that an artificial neural network with trillions of neurons can never be intelligent, no matter the structure?

Or is the claim that this particular neural network with trillions of neurons is not intelligent? If so, what is the reasoning?

> It is not smart

"Not smart" = "not able to reason intelligently".

Is that a falsifiable claim?

What would the empirical test look like that would show us if the claim is correct or not?

Look, I realize that "GPT-4 is intelligent" is an extraordinary claim that requires extraordinary evidence.

But I think we're starting to see such extraordinary evidence, illustrated by the examples below.

https://openai.com/research/gpt-4 (For instance, the "Visual inputs" section)

Microsoft AI research: Many convincing examples, summarized with:

"The central claim of our work is that GPT-4 attains a form of general intelligence, indeed showing sparks of artificial general intelligence.

This is demonstrated by its core mental capabilities (such as reasoning, creativity, and deduction), its range of topics on which it has gained expertise (such as literature, medicine, and coding), and the variety of tasks it is able to perform (e.g., playing games, using tools, explaining itself, ...)."

https://arxiv.org/abs/2303.12712

Re: Simply explained: How does GPT work?

#175
On the other hand, many people who are not ready to change, who do not have the skills or who cannot afford to reeducate are threatened.

That's me. After programming since the '80s, I'm just so tired. So much work, so much progress, so many dreams lived or shattered. Only to end up here at this strange local maximum, with so much potential, destined to forever run in place by the powers that be. The fundamentals formula for intelligence and even consciousness materializing before us as the world burns. No help coming from above, so support coming from below, surrounded by everyone who doesn't get it, who will never get it. Not utopia, not dystopia, just anhedonia as the running in place grows faster, more frantic. UBI forever on the horizon, countless elites working tirelessly to raise the retirement age, a status quo that never ceases to divide us. AI just another tool in their arsenal to other and subjugate and profit from. I wonder if a day will ever come when tech helps the people in between in a tangible way to put money in their pocket, food in their belly, time in their day - independent of their volition - for dignity and love and because it's the right thing to do. Or is it already too late? I don't even know anymore. I don't know anything anymore.

Re: Simply explained: How does GPT work?

#176
post #148

Earlier quoted context omitted.

> Sometimes for GPT to "just" complete the next word in a way that humans find plausible, it must, along the way, develop a model of the world, theory of mind, abstract reasoning. etc. I did an experiment recently where I asked ChatGPT to "tell me an idea [you] have never heard before". ChatGPT replied with what sounded like an idea for a startup, which was delivering farm-fresh vegetables to customers' doors. This i…

I think it's certainly fair to say that GPT's "reasoning" is different from human reasoning. But I think the core debate we're having is whether the difference really matters in some situations. Certainly, Midjourney's "creativity" is different from human creativity. But it is producing results that we marvel at. It's creative not because it's doing the exact same philosophical thing humans do, but because it can pro…

> It's creative not because it's doing the exact same philosophical thing humans do, but because it can produce the same effect.

Absolutely, and I hope none of my comments are taken in a way that disparages how amazing ChatGPT and Stable Diffusion et al. are. I'm just debating how humanlike they are.

> Is that really an illusion (implying its fake and potentially useless)

I don't think that because it's an illusion means that its useless. Magnets look like telekinesis, but that effect being an illusion doesn't mean that magnets are useless; far from it, and once we admit that they are what they are, they become even more useful.

> Plus, different models will excel at different thing. GPT's model will excel at synthesizing answers from far more information than a single human will ever be able to know. Does it really matter if it's not identical to human reasoning on a philosophical or biological level, if it can do things humans can't do?

It only matters if people are trying to say that ChatGPT is essentially human, that idea is all I was replying to. I completely agree with you here.

Re: Simply explained: How does GPT work?

#177
post #12

I've been using GPT4 to code and these explanations are somewhat unsatisfactory. I have seen it seemingly come up with novel solutions in a way that I can't describe in any other way than it is thinking. It's really difficult for me to imagine how such a seemingly simple predictive algorithm could lead to such complex solutions. I'm not sure even the people building these models really grasp it either.

I've started to suspect that generating code is actually one of the easier things for a predictive text completion model to achieve. Programming languages are a whole lot more structured and predictable than human language. In JavaScript the only token that ever comes after "if " is "(" for example.

> I've started to suspect that generating code is actually one of the easier things for a predictive text completion model to achieve.

> Programming languages are a whole lot more structured and predictable than human language.

> In JavaScript the only token that ever comes after "if " is "(" for example.

But isn't that like saying that it's easy to generate English text, all you need is a dictionary table where you randomly pick words?

(BTW, keep up the blog posts, I really enjoy them!)

Re: Simply explained: How does GPT work?

#178
post #148

Earlier quoted context omitted.

> Sometimes for GPT to "just" complete the next word in a way that humans find plausible, it must, along the way, develop a model of the world, theory of mind, abstract reasoning. etc. I did an experiment recently where I asked ChatGPT to "tell me an idea [you] have never heard before". ChatGPT replied with what sounded like an idea for a startup, which was delivering farm-fresh vegetables to customers' doors. This i…

I think it's certainly fair to say that GPT's "reasoning" is different from human reasoning. But I think the core debate we're having is whether the difference really matters in some situations. Certainly, Midjourney's "creativity" is different from human creativity. But it is producing results that we marvel at. It's creative not because it's doing the exact same philosophical thing humans do, but because it can pro…

Most arguments that AI can't really reason/think/invent essentially reduce to defining these terms as things only humans can do. Even if you had an LLM-based AGI that passes the Turing test 100% of the time, cures cancer, unites quantum physics with relativity, and so on, many of the people who say that ChatGPT can't reason will keep saying the same thing about the AGI.

Re: Simply explained: How does GPT work?

#179
post #160

Earlier quoted context omitted.

chatgpt doesnt just feed us back answers we already taught it. It learned relationships and semantics so it can apply that knowledge to do something novel. For instance, I took the basic of a dream and told it to turn it into a short story. The short story wasn't bad. I said make it more exciting, it updated the story such that one of the cars exploded. I guess chatgpt learned excitement from michael bay.

(I'm going to be brusque for the sake of the argument, I very much could be wrong and I don't even know how much I believe of the argument I'm making.) > chatgpt doesnt just feed us back answers we already taught it True, there is some structure to the answers we already taught it that it statistically mimics as well. > It learned relationships and semantics so it can apply that knowledge to do something novel Can yo…

I took the basics of a dream I had, and asked it to turn it into a short story. the result was pretty good. Is it using stuff already to seed its responses? sure, but thats what we do to. Nothing you do or say wasn't taught to you. But these are not simply parroting responses.

I said this to chatgpt: I had a dream that me and my friend were in a car accident, and we had a choice in deciding how to use 1 hour. we could use it to live out our last final moments or spend that hour on trying to think of a solution to survive the wreck. turn this into a short story.

As they sat trapped in their overturned car, Maria and her best friend Rachel had a dilemma. They only had one hour left to live and had to choose how to spend it. Rachel suggested they spend their final moments together, talking about all their favorite memories. But Maria had a different idea. She proposed they use the hour to think of a solution to survive the wreck. Rachel was hesitant at first, but ultimately agreed. For the next hour, they brainstormed every possible escape plan they could think of. They tried pushing against the doors, breaking the windows, and even attempting to signal for help. Just as they were about to give up hope, Maria remembered a tool in the trunk of the car. She crawled over to retrieve it and used it to pry open the door. Rachel followed her lead, and they finally escaped the car just in time. As they stumbled away from the wreck, both girls were in shock that they had made it out alive. They hugged each other tightly, grateful for the hour they spent trying to find a solution instead of giving up. From that moment on, they made a promise to never take a single moment for granted.

Re: Simply explained: How does GPT work?

#180

Earlier quoted context omitted.

We need to form some sort of guild of engineers who think Deleuze, Latour, Lacan et caterva should be read within our disciplines.

Please no! Read systems neuroscience. Like Hassabis does. Or if of a philosophical persuasion, then Dennett or Rorty.

Much of cognitive science reinvents wheels that had been established in the 1920s and 1930s already, namely in sociology of knowledge and related fields. fRMI actually often confirms what had been already observed in a psychoanalytic context. (I don't think it's a good general advice to totally ignore what is already known.)
Post reply on HN