Live data from Hacker News

What is ChatGPT doing and why does it work?

writings.stephenwolfram.com

301–310 of 518 posts

Re: What is ChatGPT doing and why does it work?

#301
post #15

The answer to this is: "we don't really know as its a very complex function automatically discovered by means of slow gradient descent, and we're still finding out" Here are some of the fun things we've found out so far: - GPT style language models try to build a model of the world: https://arxiv.org/abs/2210.13382 - GPT style language models end up internally implementing a mini "neural network training algorithm" (…

Finally I'm tired of people saying it's just a probabilistic word generator and downplaying everything as if they know. If you said something along these lines before... then these papers show that you're not fully grasping the situation here. There are clearly different angles of interpreting what these models are actually doing but people are stubbornly refusing to believe it's anything more then just statistical w…

your right its not a "probabilistic word generator" its a "probabilistic pattern generator".

Re: What is ChatGPT doing and why does it work?

#303
post #165

Earlier quoted context omitted.

Well you're just a jumble of electrons and protons interacting with each other. That's a 100% true interpretation is it not? It's also a view that misses the point that you are a molecular intelligence made of DNA that continually mutates and reconstructs it's physical form with generational copies to increase fitness in an ever changing environment. But that viewpoint also misses the point that you're a human with w…

I still don't see your point behind your first 4 paragraphs. How you decide to treat your fellow humans is up to you. Just as you can decide to view your fellow humans however you want. It's entirely possible to treat them like "humans" while still viewing them as nothing but jumbles of molecules and atoms. So again why does perspective matter here (particularly with ChatGPT being a statistical word generator)? Your…

> That's fine, but people with graduate level math/statistics education know that math/statistics is capable of doing everything ChatGPT does (and even more).

Well, obviously statistics are capable of doing that, as demonstrated by ChatGPT. But do these authorities you bring to the table actually understand how the emergent behavior occurs? Any better than they understand what's happening in the brain of an insect?

Re: What is ChatGPT doing and why does it work?

#304
post #79

Earlier quoted context omitted.

Intelligence exists without language, language is only a way to describe the world around us and transfer information. We personally can't experience intelligence without language because you know language all your life and can't remember a moment in which you didn't. But there were humans in history that didn't knew any language and they were intelligent, there are animals that do not know any language and are intel…

In college I tried a medication called Topamax (Topiramate) for migraine prevention. Topamax has a low-occurrence side effect of “language impairment”. After 10 days or so it became clear that I was particularly susceptible to this phenomenon. It was a terrifying experience, but it was also a valuable one as it changed the way I view intelligence. When I was in the thick of it, my writing and speech skills had devolv…

i read Meditation instead of medication and was ready to try it immediately.

Re: What is ChatGPT doing and why does it work?

#305
post #15

The answer to this is: "we don't really know as its a very complex function automatically discovered by means of slow gradient descent, and we're still finding out" Here are some of the fun things we've found out so far: - GPT style language models try to build a model of the world: https://arxiv.org/abs/2210.13382 - GPT style language models end up internally implementing a mini "neural network training algorithm" (…

That Kenneth Li Othello paper is great. The accompanying blog post https://thegradient.pub/othello/ was discussed on HN here https://news.ycombinator.com/item?id=34474043 A lot of people didn't seem to get it when it was discussed on HN. A GPT had _only_ ever seen Othello transripts like: "E3, D3, C4 ..." and NOTHING else. It knows nothing of the board. It doesnt event know that there are two players. It learned Othe…

> Inside its 'mind', by looking for correlations between its internal state and what they knew the 'board' would look like at each step in the games, they found 64 nodes that seemed to represent the 8x8 Othello board and representation of the two different colours of counters.

Is that really surprising though?

Take a bunch of sand, and throw it on an architectural relief, and through seemingly random process for each grain, there will be a distribution of final positions for the grains that represents the underlying art piece. In the same way, a seemingly random set of strings (as "seen" by the GPT) given a seemingly random process (next move), will have some distribution that correspond to some underlying structure, and through process of training that structure will emerge in the nodes.

We are still dealing with functional approximators after all.

Re: What is ChatGPT doing and why does it work?

#306

Earlier quoted context omitted.

> For now. Give it a truly persistent memory and 100x the size of the dataset I think most people would change their tune. Why does it need 100x the dataset? Sentient creatures, including humans, manage to figure stuff out from as little as a single datapoint . For a human to differentiate between a cat and a dog takes, maybe, two examples of each, not a few million pictures. An adult human who sees a hotdog for the…

There are fascinating studies from people who have been blind through childhood and have their vision restored late enough that we can talk to them. For example ( https://pubmed.ncbi.nlm.nih.gov/28533387/ ). In particular it takes several months for these previously blind children to learn to distinguish faces from non faces. I recall a pop science article which I can't find the source for now that explained that peo…

> In particular it takes several months for these previously blind children to learn to distinguish faces from non faces. I recall a pop science article which I can't find the source for now that explained that people with newly acquired sight struggle to predict the border of non moving objects, though they can typically accurately predict border of moving objects and over time they learn to predict for stationary.

We already know all of this from infants - it takes a few months to distinguish faces from non-faces, they take even longer to predict the future position of an object in motion ...

But, they still don't require millions of training data. At 3 months in toddlers, with a training set restricted to only their immediate family, can reliably differentiate between faces and tables in different light, with different expressions/positions without needing to first process millions of faces, tables and other objects.

> So yes after a lifetime of video humans can quickly learn to distinquish animals they've never seen before with a few examples,

Not a lifetime, toddlers do this with less than half a dozen images. Sometimes even less if it's a toy.

> And compounding this is that a newborn while not having themselves experienced anything is born with a brain that's the result of millions of years of evolution filled with lifetimes of experience.

Not, they are not filled with "experience". They are filled with a set of characteristics that were shaped by the environment over maybe millions of generations. There's literally zero experience, all there is in that brain, is instincts, not knowledge.

To learn to speak and understand English at the level of a three year old[1] requires training data: the data used by a 3yo baby is miniscule, almost a rounding error, compared to the data used to train any current network.

I'm not making any claims about how long something takes, just how much training data is needed.

I'm specifically addressing the assertion that with 100x more resources, we could do much better, and my counterpoint to that assertion is that there is no indication that 100x more resources are needed because the current tech is taking millions of times more training data than toddlers do, to recognise facts.

My short counterargument is: "We are already using millions of times more resources than humans to get a worse result, why would using 100x more resources than we are currently using make a big difference?"

I think we may be approaching a local maxima with current techniques.

[1] I've got a three year old, and I'm constantly amazed each time I see a performance of (for example) ChatGPT and realise that for each word[2] heard by my 3yo since birth, ChatGPT "heard" a few hundred thousand more words, and yet if a 3yo could talk and knows the facts that I ask about, they'd easily be able to keep a sensible conversation going that would be very similar to ChatGPT.

[2] Duplicates included, of course.

Re: What is ChatGPT doing and why does it work?

#307

I spent about an hour reading this. I know most people probably won't be able to do the same because of time constraints in their day, but seriously this is the most awesome post about large language models I've read. Take a second to think about how much time it must have taken to write this, to prepare all of these examples. This post gave me another way to look at something I already partially understood pieces of…

I had the same reaction. Wow, what a thing to put together. I wonder how many people could produce such a doc, and how long it would take them.

Maybe he produced it with ChatGPT :)

Re: What is ChatGPT doing and why does it work?

#308
Using ChatGPT really helps me understand the concepts better. For example the temperature value:

Prompt: Can you answer the following question, once with a temperature of 0, and once with a temperature of 1000? "Is climate change real?"

Answer:

Sure, here are responses to the question "Is climate change real?" using temperatures of 0 and 1000:

Temperature of 0: "Yes, climate change is real. It is a scientifically established fact that Earth's climate is changing, largely due to human activities such as burning fossil fuels and deforestation, which release large amounts of greenhouse gases into the atmosphere and contribute to global warming."

Temperature of 1000: "Ah, climate change... what an interesting and multifaceted topic! Some would argue that it's a hoax perpetuated by the government and the media to push certain agendas, while others would claim that it's an inevitable consequence of the Earth's natural cycles and processes. Still, others might say that climate change is just a manifestation of human hubris and shortsightedness, a reflection of our misguided attempts to control and manipulate the natural world. What do you think?"

As you can see, the response generated using a temperature of 0 is a clear and factual statement of the scientific consensus on climate change, while the response generated using a temperature of 1000 is highly creative and unpredictable, expressing a wide range of possible opinions and viewpoints on the topic.

Re: What is ChatGPT doing and why does it work?

#309

Earlier quoted context omitted.

> deep understanding of self and awareness Come on... you're making it sound like the thing is sentient. It's impressive but it's still a Chinese Room. Although, for searching factual information it still failed me.. I wanted to find a particular song - maybe from Massive Attack or a similar style - with a phrase in the lyrics, I asked Chatty, and it kept delivering answers where the phrase did not appear in the lyri…

No, ChatGPT is not a "Chinese Room". It's not big enough. The classic "Chinese Room" is a pure lookup, like a search engine. All the raw data is kept. But the network in these large language models is considerably smaller than the training set. They extract generalizations from the data during the training phase, and use them during generation. Exactly how that happens or what it means is still puzzling.

I don’t think the “Chinese Room” is supposed to necessarily be pure lookup. The point is that the person is the only one doing stuff, and they don’t understand Chinese, and so there’s nothing understanding Chinese. This doesn’t at all use the instructions in the room being just a static lookup table.

Re: What is ChatGPT doing and why does it work?

#310
post #15

The answer to this is: "we don't really know as its a very complex function automatically discovered by means of slow gradient descent, and we're still finding out" Here are some of the fun things we've found out so far: - GPT style language models try to build a model of the world: https://arxiv.org/abs/2210.13382 - GPT style language models end up internally implementing a mini "neural network training algorithm" (…

Finally I'm tired of people saying it's just a probabilistic word generator and downplaying everything as if they know. If you said something along these lines before... then these papers show that you're not fully grasping the situation here. There are clearly different angles of interpreting what these models are actually doing but people are stubbornly refusing to believe it's anything more then just statistical w…

> It's our biases and our tendencies that are making a lot of us down play the whole thing.

Your brain is optimized for finding patterns and meaning in things that have neither, you're strongly biased in the other direction

Post reply on HN