Live data from Hacker News

What is ChatGPT doing and why does it work?

writings.stephenwolfram.com

331–340 of 518 posts

Re: What is ChatGPT doing and why does it work?

#331

Wow this is 19,000 words. I like his summary at the end: At some level it’s a great example of the fundamental scientific fact that large numbers of simple computational elements can do remarkable and unexpected things. And this: ... But it’s amazing how human-like the results are. And as I’ve discussed, this suggests something that’s at least scientifically very important: that human language (and the patterns of th…

I don’t think the language point is particularly revelatory - we’ve lived with quite effective machine translation for a long while now. But it’s certainly unexpected that large swathes of complex knowledge can be gathered and represented this way (as patterns of patterns of patterns). Consequentially, ChatGPT is a still a fairly uninteresting pattern matching machine in itself. It has very static knowledge and no way to reason or ponder or evaluate or experiment between that knowledge and the world beyond, as anyone trying to use ChatGPT to get ‘correct’ answers and not just vaguely cromulent ideas is finding. We’ve perhaps proven that machines can know what we can know, but can’t think as we can think. I would not bet against the latter being solved in my lifetime though.

Re: What is ChatGPT doing and why does it work?

#332

Earlier quoted context omitted.

> For now. Give it a truly persistent memory and 100x the size of the dataset I think most people would change their tune. Why does it need 100x the dataset? Sentient creatures, including humans, manage to figure stuff out from as little as a single datapoint . For a human to differentiate between a cat and a dog takes, maybe, two examples of each, not a few million pictures. An adult human who sees a hotdog for the…

>Why does it need 100x the dataset? Sentient creatures, including humans, manage to figure stuff out from as little as a single datapoint. Human brains are not quite blank slates at birth. They're predisposed to interpret and quickly learn from the sort of inputs that their ancestors were exposed to. That is to say, the brain, which learns, is also the result of a learning process. If a mad scientist rewired your bra…

This.

Also consider that a human brain that is able to figure stuff from as little as a single datapoint is normally exposed to at least 4 years of massive and socially "directed" multimodal data patterns.

As many cases of feral childs have shown, those humans not "trained" in their first years of life will never be able to harness language and therefore will never be able to display human-level intelligence.

Re: What is ChatGPT doing and why does it work?

#334

Earlier quoted context omitted.

I was responding to someone claiming humans learn these things with only one or two examples. I am aware of that GPT3 pretty much scraped every bit of text Open AI could find on the internet and I agree that probably makes it less example efficient than humans. But I also think this critique is slightly unfair, your brain has had the benefit of thousands of lifetimes of experience informing their structure and in bui…

The human brain hasn't had to "evolve" to learn writing. Our brain hasn't really changed for many thousands of years and writing has only been around for about 5000 years so we can't use the argument that "human brains have evolved over millions of years to do this" - it's not true. GPT3 essentially needs millions of human years of data to be able to speak English correctly but still make obvious mistakes to us, so t…

Writing was specifically designed (by human brains) to be efficiently learnable by human brains.

Same for many other human skills, like speaking English, that we expect GPT to learn.

Re: What is ChatGPT doing and why does it work?

#335

This misses that key point that all this prediction can give rise to what looks like astonishing human-level creativity and across many genres. The last decade and half have shown us that with enough data we can pick out patterns well enough to be able to "categorize". But to create , that seemed like a whole another human level outside the realm of mere prediction. Turns out it isn't. What exactly allows LLMs to hav…

I'd argue that the reason you're probably impressed with this is because it's outside your domain. You are probably less than impressed that it can generate entire code segments, or fluently answers arbitrary questions because it's easy to see how a training on Stack Exchange and other such sources can easily generate such output. By contrast somebody into rap, but with little knowledge of programming, would probably…

That could be. As a programmer I am still very impressed by its ability to create a chrome extension for a very specific use-case that I imagine there is not much data on. This rap battle seem like a very similar cross-genre fusion. It would be interesting to see some study on how well LLMs do transfer learning. Do the meta-patterns learned from code segments (what Wolfram refers to as semantic grammar) be used to pick up meta-patterns from a completely different genre (rap battles, medical literature, legalese etc.) with very few examples. It does seem like it is doing such transfer learning, but yes tough to say if there's say without knowing what data it actually had access to. Seeing an open-source replication that also analyzes how well it does across genres and training data size from that genre would be nice.

Re: What is ChatGPT doing and why does it work?

#336

Earlier quoted context omitted.

Ahh, so being as intelligent as an average intelligence is no longer sufficient to declare it intelligent. Now it must surpass all of our achievements. 99.9% of people will never "create some substantiative achievement".

I said to call it super-intelligent. To demonstrate super-intelligence it would need to demonstrate real creative powers that are beyond us in both scope and direction. That isn't necessary to prove that this is productive work; but I think it is necessary to temper some of the enthusiasm I see in this thread that all but call it a super-intelligence.

Maybe you are right.

As an observation: a human of normal intelligence but with much better access to a calculator and to Wikipedia, or even just external storage (faster than pen-and-paper) would already be super-human.

Re: What is ChatGPT doing and why does it work?

#337
post #15

The answer to this is: "we don't really know as its a very complex function automatically discovered by means of slow gradient descent, and we're still finding out" Here are some of the fun things we've found out so far: - GPT style language models try to build a model of the world: https://arxiv.org/abs/2210.13382 - GPT style language models end up internally implementing a mini "neural network training algorithm" (…

Such a hacker news comment. The title isn't a question Stephen Wolfram is asking you a question, it's the title of an article he's written that answers the question.

Re: What is ChatGPT doing and why does it work?

#338
post #165

Earlier quoted context omitted.

I still don't see your point behind your first 4 paragraphs. How you decide to treat your fellow humans is up to you. Just as you can decide to view your fellow humans however you want. It's entirely possible to treat them like "humans" while still viewing them as nothing but jumbles of molecules and atoms. So again why does perspective matter here (particularly with ChatGPT being a statistical word generator)? Your…

Ok let me make this more clear. I choose how to view things, yes this is true. But if I choose to treat human beings as jumbles of molecules, most people would consider that viewpoint flawed, inaccurate and slightly insane. Other humans would think that I'm in denial about some really obvious macro effects of configuring molecules in a way such that it forms a human. I can certainly choose to view things this way, bu…

The paper is discussed by another commenter which brings interesting points how that model is built and it may not be what you think.

I agree with what you said regarding how we choose to view things, but I think you also have a bias/belief that you want it to be something more, instead of being more neutral and scientific: we know the building blocks, we have to study the emerging behaviors, we can’t assume the conclusion. One paper is not enough, we have to stay open.

Re: What is ChatGPT doing and why does it work?

#339

Earlier quoted context omitted.

It's clearly early technology, so it's not perfect. But what it is able to get right is clear proof it's more then what you think: https://www.engraved.blog/building-a-virtual-machine-inside/ Read to the end. The beginning and middle doesn't show off anything too impressive. It's the very end where chatGPT displays a sort of self awareness. Also here's a scientific paper showing that LLMs are more then a chinese room…

I asked Google Home what the definition of self-awareness is, and it says "conscious knowledge of one's character s and feelings.". But me saying "ChatGPT surely doesn't have feelings, so it can't be self-aware!" would be a simple cop-out/gotcha response. I guess it's a Chinese Room, that when you ask about Chinese Rooms, can tell you what those things are. I almost said the word "aware" there, but the person in the…

Therapy sometimes uses a method called exposition. E.g. if one has an irrational fear of elevators, they can gradually expose themselves to it. Stand before it then leave. Call it and look inside. Enter it on the first floor and exit without riding. After few weeks or months they can start using it, because the fear response reduces to manageable levels. Because nothing bad happens (feedback).

One may condition themselves this way to torture screams, deaths, etc. Or train scared animals that it’s okay to leave their safe corner.

And nothing happens to you in a seismically inactive areas when an earthquake ruins whole cities somewhere. These news may touch other (real) fears about your relatives well-being, but in general feeling sad for someone unknown out there is not healthy even from the pov of being a biologically human (watch emphasis, the goal isn’t to play cynic here). It’s ethical, humane, but not rational. The same amount of people die and become homeless every year.

What I’m trying to say here is: feelings are our builtin low-cost shortcut to thinking. Feelings cannot be used as a line that separates conscious from non-conscious or non-self-aware. The whole question “is it c. and s.a.?” refers completely to ethics, which are also our-type-of-mind specific.

We may claim what Chinese Room is or isn’t, but only to calm ourselves down. But in general it’s just a type of consciousness, one of a relatively infinite set. We can only decide if it’s self-ethical to think about it in some way.

Re: What is ChatGPT doing and why does it work?

#340
post #188

Earlier quoted context omitted.

For now. Give it a truly persistent memory and 100x the size of the dataset I think most people would change their tune.

Hasn't it already been trained on what is effectively the entire contents of the scrapable internet? There isn't another 10x to be had there, let alone 100x. I assume that whatever future improvements we get from improving algorithms (or perhaps through throwing more compute at it), not through larger datasets.

There might not be another 100x of written language.

But we noticed that training your neural networks on multiple tasks actually works well. So we could start feeding our models eg audio and video.

With lots of webcams we can make arbitrary amounts of new video footage. That would also allow the language model to be grounded more in our 3d reality.

(Granted, we only know as a general observation that training the same network for multiple tasks 'forces' that network to become better and abstract and generalise. Nobody has yet publicly demonstrated an application of that observation to training language models + video models.)

Another avenue: at the moment those large language models only see each example once, if I remember right. We still have lots of techniques for augmenting training data (eg via noise and dropout etc), or even just presenting the same data multiple times without overfitting.

Post reply on HN