Earlier quoted context omitted.
You haven't specified what model did you use, and the green ChatGPT icon in the shared conversation usually signifies GPT-3.5 model. Here's my attempt at similar conversation — it seems GPT-4 is able to visualise the board and at least do a valid first move. https://chat.openai.com/share/98427e21-678c-4290-aa8f-da8e93...
Interesting. The model was whatever was up that that time, so probably was 3.5 if you say so.
Large language models lack deep insights or a theory of mind
111–120 of 270 posts
Re: Large language models lack deep insights or a theory of mind
#112For me, the entire AGI conversation is hyperbolic / hype. How can we infer intelligence to something when we, ourselves, have such a poor (none) grasp of what makes us conscience? I'm associating intelligence with consciousness - because it seems correlated. Are we really ready to associate "AGI" with solving math problems ("new Q algo.")? That seems incredibly naive & reinforces my opinion that LLM's are much more l…
Couldn't agree more. How about this -- I think we've already reached AGI. Let me know if this tracks: Pick a set of tasks that can be considered AGI tasks. Provided the task sequences can be compared as closer to AGI or further from AGI, we can create a reward model using the same techniques as were used by ChatGPT via RLHF. Thus, for any definition of AGI that is meaningful and selectable, even if subjectively selec…
Let's say I'm deep in a coding problem. A co-worker comes by and says "How did your team do in the game yesterday?". I say, "Um, uh... sorry, my head's not there right now." It takes us time to swap between mental "spaces".
So, if I have an AGI (defined as having a trained model for almost everything, even if that turns out to be a large number of different models), if it has to load the right model before it can reason on that topic, then that's pretty human-like. (As long as it can figure out which model to load...)
The one thing missing is that (at least some) humans can figure out linkages between different mental "spaces", to turn them into a more coherent holistic mental space, even if they don't have (all of) each space at front-of-mind at any moment. I'm not sure if this flavor of an AGI could do that - could see the connections between different models.
Re: Large language models lack deep insights or a theory of mind
#113Earlier quoted context omitted.
Our tests against current cars found that they didn't perform well on transatlantic flights... But who knows what the future holds? Maybe we should test them again next year. LLM names an specific product, aimed at solving an specific problem.
our tests of combustion engine driven crankshafts connected to a spinning mechanism can perform well in both transatlantic flights and cross country road trips.
Re: Large language models lack deep insights or a theory of mind
#114I think that if they would, that would be very surprising and indicative of a lot of wastefulness inside the model architecture. All these tests are simple single prompt experiments, so the LLM's get no chance to reason about their responses. They're just system 1 thinking, the equivalent of putting a gun to someone's head and asking them to solve a large division in 2 seconds. I bet a lot of these experiments would…
The equivalent for a human would be an reflexive response to a question, the kind you could immediately answer after being woken up at 3am in the morning. That type of answer has been deeply trained into the human networks and also requires no deep insight. But if a human is allowed time and internal reasoning iterations, so should the LLM when determining if it has deep insight. Right now we're simply observing inpu…
From a computer science point of view: a single prompt/response cycle from a LLM is equivalent to a pure function; the answer is a function of the prompt and the model weights and is fundamentally reducible to solving a big math equation (in which each model parameter is a term.)
It seems almost self evident that "reasoning" worthy of the name would involve some sort of iterative/recursive search process, invoking the model and storing/reflecting/improving on answers methodically.
There's been a lot of movement in this direction with tree-of-thought/chain-of-thought/graph-of-thought prompting, and I would bet that if/when we get AGI, it's a result of getting the right recursive prompting pattern + retrieval patterns + ensemble models figured out, not just making ever-more-powerful transformer models (thought that would certainly play a role too.)
The LLM isn't the whole brain. Just the area responsible for language and cultural memory.
Re: Large language models lack deep insights or a theory of mind
#115Earlier quoted context omitted.
This reminds me of The Last Question by Isaac Asimov. I also think if we stopped expecting all LLMs to have an immediate answer, it would be relatively easy to shim some kind of "conscience" to direct the output in different ways. Similar to the safeties already in place in LLMs, but instead of it just saying "NO DON'T SAY THAT" it can dialog internally to change what the output is until it reaches what it believes t…
It would have an emotional reaction to certain "thought constructs" and would be guided by that. Or we could just give them three laws
Otherwise the laws will have to be implemented as weights or during training so the model explicitly knows the laws and would never even be capable of doing something against them.
Re: Large language models lack deep insights or a theory of mind
#116Earlier quoted context omitted.
I think they may be referring to the principle task that consciousness serves in humans, which is to rationalize decisions we've already made subconsciously to other people so they will help us. The conscious "why" comes after the decision. In that sense it's exactly the kind of bullshit machine that LLMs are.
A thought experiment: what kind of functional MRI result would convince you that human consciousness is real and an important part of decision making? Note: if the result is someone reporting having made a decision before brain activity is seen, my next question is going to be "How does that work?"
Here's a study from 2009 showing that the problem gets solved but it takes up to 8 seconds for the person to announce that the problem is solved.
Showing that even for problem solving, there's a giant unconscious machinery that is active. Although I assume they cannot discount the effects of the conscious bits (struggle, suffering, anxiety, active as a mirror of that problem solving effort being a feedback loop to this unconscious effort).
Re: Large language models lack deep insights or a theory of mind
#117The eval is a weird, noisy visual task (picture of astronaut with “care packages”). Their results are hopelessly narrow.
A better eval is to use actual scientifically tested psychology test on text (the native and strongest domain for LLMs), for example the sort of scenarios used to gauge when children develop theory of mind (“Alice puts her keys on the table then leaves the room. Bob moves the keys to the drawer. Alice returns. Where does she think the keys are?”) which GPT-4 can handle easily; it is very clear from this that GPT has a theory of mind.
A negative result doesn’t disprove capabilities; it could easily show your eval is garbage. Showing a robust positive capability is a more robust result.
Re: Large language models lack deep insights or a theory of mind
#118In Buddhism there’s the idea that our core self is awareness, which is silent - it doesn’t think in a perceptible way, it doesn’t feel in a visceral way, but it underpins thought and feeling, and is greatly impacted by it. A large part of meditation and “release of suffering” is learning to let your awareness lead your thinking rather than your thinking lead your awareness. To be clear, I think this is in fact a corr…
https://en.wikipedia.org/wiki/Moravec%27s_paradox While you're adding a bunch of eastern philosophy to it, we need to take a step back from 'human' intelligence and go to animal and plant intelligence to get a better idea of the massive variation in what covers thought. In animal/insects we can see that thinking is not some binary function of on or off. It is an immense range of different electrical and chemical proc…
Is this actually true? I thought it just involved a different part of the brain. Is there actually no brain involvement? Sure it does not need your awareness or decision making, but no brain? I find that hard to believe.
Re: Large language models lack deep insights or a theory of mind
#119Earlier quoted context omitted.
> Maybe the soul is social My pet theory about human consciousness is that is that consciousness is simply recursive theory of mind. Theory of mind [1] is our ability to simulate and reason about the mental states of others. It's how we predict what people are thinking and how they will react to our actions, which is critical for choosing how to act in a social environment. But when you're thinking about what's in so…
Theory of mind is interesting but one wouldn't want to hinge consciousness upon it. That direction would likely contain weird outcomes if the science progressed, something like "Dogs are barely-conscious due to their pack structure, they have a couple levels of recursive theory of mind but they can't sustain it as deep as we can. But cats didn't have that pack structure, they're not conscious at all." Or, "this perso…
Let's say the soul is exclusively located in eye-to-eye contact. Theres a lot of information in how that contact is broken, how long its broken for, and what happens in between.
(Enemy's-gate-is-down-style reorientation)
Re: Large language models lack deep insights or a theory of mind
#120This is a terrible eval. Do not update your beliefs on whether LLMs have Theory of Mind based on this paper. The eval is a weird, noisy visual task (picture of astronaut with “care packages”). Their results are hopelessly narrow. A better eval is to use actual scientifically tested psychology test on text (the native and strongest domain for LLMs), for example the sort of scenarios used to gauge when children develop…
Aren't you confusing having a theory of mind with being able to output the right answer to a test? Isn't your proposed evaluation especially problematic because an "actual scientifically tested psychology test" is likely in the training data along with a lot of discussion and analysis of that test and the correct and incorrect answers that can be given?