Live data from Hacker News

Simply explained: How does GPT work?

confusedbit.dev

291–300 of 392 posts

Re: Simply explained: How does GPT work?

#291
post #160

Earlier quoted context omitted.

chatgpt doesnt just feed us back answers we already taught it. It learned relationships and semantics so it can apply that knowledge to do something novel. For instance, I took the basic of a dream and told it to turn it into a short story. The short story wasn't bad. I said make it more exciting, it updated the story such that one of the cars exploded. I guess chatgpt learned excitement from michael bay.

(I'm going to be brusque for the sake of the argument, I very much could be wrong and I don't even know how much I believe of the argument I'm making.) > chatgpt doesnt just feed us back answers we already taught it True, there is some structure to the answers we already taught it that it statistically mimics as well. > It learned relationships and semantics so it can apply that knowledge to do something novel Can yo…

Sorry about this, but I couldn't resist:

GTP4, rewrite the above message to be less brusque

I hope you don't mind me sharing a different perspective for the sake of discussion. I might be mistaken, and I'm not entirely sure how much I believe in the argument I'm presenting.

It's true that ChatGPT doesn't only provide answers based on what it has been taught, but it also uses the structure of those answers to create statistically similar responses. However, when it comes to demonstrating novelty, I think we might be underestimating the vast amount of information and variety that humans have written about and shared online. While anything we ask ChatGPT to do might be new to us personally, it's highly likely that it has already been thought of and documented online, and ChatGPT is simply providing a similar response based on its prior knowledge.

This phenomenon, where ChatGPT has significantly more training data and experience than any single human, leading to the illusion of originality, is quite intriguing. For instance, when it comes to image generation, we might experience a period of "artistic inbreeding" because we, as individuals, are not aware of everything others have done. We may perceive something like Midjourney's output as moving and original, when in reality, it could just be a slight variation of someone else's work that we haven't seen before.

Please don't take this as me being confrontational; I genuinely respect your opinion and experiences, and I'm enjoying our conversation.

Re: Simply explained: How does GPT work?

#292

Earlier quoted context omitted.

> I would already, without hesitation, describe GPT4 as strictly more intelligent than my cat Well if we're going to define intelligence based one what you believe it is then why don't you explain it? I'm not the one claiming to know what intelligence is or that we can even simulate a system capable of emulating this characteristic. So if you hold the specification for human thought I think you ought to share it with…

You're the one claiming that GPT is not in any sense, shape, or form intelligent. Such claim inevitably carries a very strong implication that you know what intelligence is.

One doesn’t have to know how thoughts are formed to have good theories and reasonable hypothesis.

Science makes progress with imperfect information all the time, including incomplete models of neurological phenomenon, intelligence, and consciousness.

Re: Simply explained: How does GPT work?

#293

Earlier quoted context omitted.

That's highly reductive of our capacities. We are not weighted transformers that can be explained in an arxiv paper. GPT, at the end of the day, is a statistical inference model. That's it. It's not going to wake up one day, decide it prefers eggs benny and has had enough of your idle chatter because of that sarcastic remark you made last week. Could we simulate a plausibly realistic human brain on silicon someday? I…

I read this and can't help but chuckle... To say that we are nowhere being able to have AGI is quite a bold statement. It was after all only a few months ago where many people also believed we were a long way away from ChatGPT-4. The confidence with which you think we are not weighted transformers or statistical inference models is also puzzling. How could you possibly know that? How do you know that that's not preci…

Ah yes, the old: you can’t prove my deity doesn’t exist argument.

Puzzling that I don’t share your faith or point of view? Why?

The point is to not ascribe properties attributed to a thing we know doesn’t have them. We can teach people how ChatGPT works without getting into pseudo-philosophical babble about what consciousness is and whether humans can be accurately simulated by an LLM with enough parameters.

Re: Simply explained: How does GPT work?

#294
post #290

Earlier quoted context omitted.

It can pass tests and exams with answers that were not included in its training corpus. For example, it passed the 2023 unified bar exam, though its training cut off in 2021. Yes, it can look at previous test questions and answers, just like human law students can. Are you therefore claiming that human law students don't engage in abstract reasoning when they take the bar exam, since they studied with tests from prev…

It is a large language model. It manipulates text based on context and the imprint of its vast training. You are not able to articulate a theory of reasoning. You are just pointing to the output of an algorithm and saying "this must mean something!" There isn't even a working model of reasoning here, it's just a human being impressed that a tool for manipulating symbols is able to manipulate symbols after training it…

It's not clear to me what point you're trying to make. Why do we need an "articulated theory of abstract reasoning" to say that passing the bar exam or writing code for novel, nontrivial tasks requires reasoning? Seems rather obvious.

Re: Simply explained: How does GPT work?

#295
post #20

Is it possible that we don’t truly know how it works? That there is some emergent behavior inside these models that we’ve created but not yet properly described? I’ve read a few of these articles but I’m still not completely satisfied.

I hate being the bearish guy during the hype cycle, but I think a lot of that is just anthropomorphizing it. They fed it TBs of human text, it spits out human text, we think it's humanesque. Of course maybe I'm wrong and it's AGI and it will find this comment and torture me for for insulting it's intelligence.

It's more like you feed a million cows into a meat grinder, then into a sausage machine, and then weirdly what appears to be a mooing cow comes out the other end.

It's weird it works when you know how it works.

Re: Simply explained: How does GPT work?

#296
post #225

Earlier quoted context omitted.

> It's a very large black box. It was trained on guessing the next word. Does that fact alone prove that it cannot have evolved certain internal structures during the training? Yes. There is interesting work to formalize these black boxes to be able to connect what was generated back to its inputs. There’s no need to ascribe any belief that they can evolve, modify themselves, or spontaneously develop intelligence. As…

> There’s no need to ascribe any belief that they can evolve, modify themselves, or spontaneously develop intelligence. But neural networks clearly evolve and are modified during training. Otherwise they would never get any better than a random collection of weights and biases, right? Is the claim then that an artificial neural network can never be trained in such a way that it will exhibit intelligent behavior? >> D…

> Is the claim then that an artificial neural network can never be trained in such a way that it will exhibit intelligent behavior?

I think it’s not likely a NN can be trained to exhibit any kind of autonomous intelligence.

Science has good models and theories of what intelligence is, what constitutes consciousness, and these models are continuing to evolve based on what we find in nature.

I don’t doubt that we can train NN, RNN, and deep learning NN to specific tasks that plausibly emulate or exceed human abilities.

That we have these deep learning systems that can learn supervised and unsupervised is super cool. And again, fully explainable maths that anyone with enough education and patience can understand.

I’m interested in seeing some of these algorithms formalized and maybe even adding automated theorem proving capabilities to them in the future.

But in none of these cases do I believe these systems are intelligent, conscious, or capable of autonomous thought like any organism or system we know of. They’re just programs we can execute on a computer that perform a particular task we designed them to perform.

Yes, it can generate some impressive pictures and text. It can be useful for all kinds of applications. But it’s not a living, breathing, thinking, autonomous organism. It’s a program that generates a bunch of numbers and strings.

But when popular media starts calling ChatGPT “intelligent,” we’re performing a mental leap here that also absolves the people employing LLM’s from responsibility for how they’re used.

ChatGPT isn’t going to I take your job. Capitalists who don’t want to pay people to do work are going to lay off workers and not replace them because the few workers that remain can do more of the work with ChatGPT.

Society isn’t threatened by ChatGPT becoming self aware and deciding it hates humans. It cannot even decide such things. It is threatened by scammers who have a tool that can generate lots of plausible sounding social media accounts to make a fake application for a credit card or to socially engineer a call centre rep into divulging secrets.

Re: Simply explained: How does GPT work?

#297
post #290

Earlier quoted context omitted.

It is a large language model. It manipulates text based on context and the imprint of its vast training. You are not able to articulate a theory of reasoning. You are just pointing to the output of an algorithm and saying "this must mean something!" There isn't even a working model of reasoning here, it's just a human being impressed that a tool for manipulating symbols is able to manipulate symbols after training it…

It's not clear to me what point you're trying to make. Why do we need an "articulated theory of abstract reasoning" to say that passing the bar exam or writing code for novel, nontrivial tasks requires reasoning? Seems rather obvious.

You are making a claim that there is some attribute of importance. For that claim to be persuasive, it should be supported with an explanation of what that attribute is and is not, and evidence for or against the meeting of those criteria. So far all you have done is say "Look at the text it puts out, isn't that something?"

It's just empty excitement, not a well-reasoned argument.

Re: Simply explained: How does GPT work?

#298
post #264

Earlier quoted context omitted.

If you read some of the studies of these new LLMs you'll find pretty compelling evidence that they do have a world model. They still get things wrong but they can also correctly identify relationships and real world concepts with startling accuracy.

No, they don't. They fail at the arithmetics ffs.

It fails at _some_ arithmetic. Humans also fail at arithmetic...

In any case, is that the defining characteristic of having a good enough "world model"? What distinguishes your ability understand the world vs. an LLM? From my perspective, you would prove it by explaining it to me, in much the same way an LLM could.

Re: Simply explained: How does GPT work?

#299
post #283
post #234

Earlier quoted context omitted.

Thanks. To get to what I think is the core of your argument (?) > ChatGPT simply "finds" training data where someone asked a similar question, and produces the likely response, which is an idea that it has actually "heard," or seen in its training data, before. I can definitely see a scenario where we manage to build an ultra-intelligent machine that can figure out any logical puzzle we put to it, but where it still…

That does seem really impressive. But don't you think that it's pretty likely that this, or something phrased slightly differently, appeared in the training data?

> But don't you think that it's pretty likely that this, or something phrased slightly differently, appeared in the training data?

I don't think so, but I could be wrong. It's definitely not "likely", see the math below.

I base that on the fact that people seemed to spend quite a bit of time trying to find the phrase "the confetti has left the cannon" that GPT-4 phrased. It seems Google search has no records of it before then?

I've seen many other examples where GPT-4 can translate sentences between using different types of idioms, and I just can't picture all these weird examples already being present on the Internet?

Do you think GPT-4 is a stochastic parrot that just has a large database of responses?

If so, how would we test that claim? What logical and reasoning problems can we give it where it fails to answer, but a human doesn't?

My understanding is that even with an extremely limited vocabulary of 32 words, you quickly run out of atoms in the universe (10^80) if you string more than 50 words together. If your vocabulary instead is 10k words, you reach 10^80 combinations after 20 words.

By training the LLMs on "fill in the missing word", they were forced to evolve ever more sophisticated algorithms.

If you look at the performance over the last 5 years of increasingly larger LLMs, there was a hockey-stick jump in performance 1-2 years ago. My hunch is that is when they started evolving structures to generate better responses by using logic and reasoning instead of lookup tables.

Re: Simply explained: How does GPT work?

#300
post #288

Earlier quoted context omitted.

Why not both? Things like philosophy or metapsychology tend to be prismatic, each framework comes with advantages and disadvantages and boundaries of its own. (A turn towards the dogmatic is something I'm pretty much expecting from the current launch of AI anyway, simply, because the productions systematically favor the semantic center. So it may be worth putting some generality against this, rather than being overly…

Lol well to answer your question literally, I think integrationist linguistics and Wittgenstein's thoughts about language use as a social action are way more relevant to understanding what's happening with LLMs (and people's naive reactions to them) than what was suggested previously as background reading.

Mind that we're are not, by any means, at any state of social interaction with LLMs. (Any such thing would be a mere hallucination on the user's side.) However, these are semantic fields, with whatever consequence comes with this. (So there may have been something said on this already, in what was known as the linguistic turn.)
Post reply on HN