Live data from Hacker News

Do LLMs pass the mirror test?

blog.pascalschuster.de

41–50 of 72 posts

Re: Do LLMs pass the mirror test?

#41
post #40

Earlier quoted context omitted.

This is just a misconception of how LLMs work and also what reasoning is. “There cannot be any reasoning embedded in the model” a strong statement, what do you mean by reasoning because by any reasonable definition I’m aware of, they clearly are able to exhibit reasoning. The fact that the pre training objective is next token loss has nothing to do with capabilities or their ability to reason. To be highly successful…

LLM output produces the illusion of reasoning. The underlying computation, however, is not reasoning.

If you don’t mind actually taking a few more words to be more specific that would be helpful because what you’re saying doesn’t really make sense at all. You don’t need to trust that the reasoning traces are all faithful representation of an internal reasoning trace. Plenty of other ways to probe models (see anthropics work using circuit tracing).

Re: Do LLMs pass the mirror test?

#42
post #40

Earlier quoted context omitted.

This is just a misconception of how LLMs work and also what reasoning is. “There cannot be any reasoning embedded in the model” a strong statement, what do you mean by reasoning because by any reasonable definition I’m aware of, they clearly are able to exhibit reasoning. The fact that the pre training objective is next token loss has nothing to do with capabilities or their ability to reason. To be highly successful…

LLM output produces the illusion of reasoning. The underlying computation, however, is not reasoning.

How do you define reasoning? What does a system have to functionally do in order to qualify for it?

Re: Do LLMs pass the mirror test?

#44
post #40

Earlier quoted context omitted.

LLM output produces the illusion of reasoning. The underlying computation, however, is not reasoning.

If you don’t mind actually taking a few more words to be more specific that would be helpful because what you’re saying doesn’t really make sense at all. You don’t need to trust that the reasoning traces are all faithful representation of an internal reasoning trace. Plenty of other ways to probe models (see anthropics work using circuit tracing).

What else is there to say? LLMs can at most regurgitate approximations of human reasoning steps in the limited forms in which they may be expressed in the training data or interpolations thereof. That's the core essence of what they are. There is no proper reasoning to be found.

Re: Do LLMs pass the mirror test?

#45

Anthropic has some mechanistic interpretabilty research on this actually. https://www.anthropic.com/research/introspection TLDR; Part 1: Testing introspection with concept injection First they find neural activity patterns they attribute to certain concepts by recording the model’s activations in specific contexts (so for example, they find the concept of "ALL CAPS" or "dogs"). Then they inject these patterns into th…

Yup, those are among the papers I was referring to in the opening parts of the piece! The difference between them and my small tests is that they all explicitly prompt the model to introspect, while I specifically didn't and kept the context perfectly "normal conversation"-shaped (minus the complete corruption of the model's outputs, of course).

There's another one that intrigued me greatly when i read about it years back. This was back when GPT-3 was state of the art. I had a lot of trouble finding it again but i did!

It's not an exact fit because the output is that of a tool rather than the model itself (though i don't think much would change if we had the model perform the arithmetic itself but altered answers similarly), but it was the first time I began to realize that just like the brain, these models have an expectation of reality that they work around. They don't necessarily 'trust' an output if it diverges significantly from this 'reality'. And that this disregard may be silent indeed (no reasoning or chain of thought here).

GPT-3 will ignore tools when it disagrees with them - https://vgel.me/posts/tools-not-needed/

Re: Do LLMs pass the mirror test?

#46

  >> Wait, looking at the prompt history, the model had a strange quirk.

  Throughout every prior thinking trace in the conversations (and, honestly, every other thinking trace across all other conversations I've had with it), the frame is always in first-person, including the moment in this one where it "noticed" the corruption: "I noticed," "I had some weird typos," "did I do that on purpose?" And then the moment the anomaly couldn't be reconciled with the self-model, the language shifted to third person: "The model had a strange quirk." Effectively, the thing doing the thinking dissociated from the thing that produced the anomalous output, as if they were two entirely different layers of the process, much in the same way a person might fumble an easy sentence and then go for something like "my brain just did something weird." Except, of course, that "me" vs "my brain" is a distinction without a difference in much the same way Gemma's "I" vs "the model" is. Gemma is the model, just as much as we are our brains.
I'll leave aside the claim that "we are our brains" - this actually reads to me like Gemma might have briefly responded as if its history came from another LLM agent and it was the next line in the chain. OTOH it might have been reading its RLHF notes a little too closely. The stuff about "my brain did X" is too anthropomorphic for my taste.

Likewise with Claude referring to "the model" - that quote sounds like something an Anthropic worker would say. Seems like a pithy little line Claude could have learned "on the job."

Re: Do LLMs pass the mirror test?

#47
post #40

Earlier quoted context omitted.

LLM output produces the illusion of reasoning. The underlying computation, however, is not reasoning.

How do you define reasoning? What does a system have to functionally do in order to qualify for it?

Reasoning includes things like proper use of logic. LLMs have been repeatedly shown to fail horribly at this.

They consistently fail at drawing basic logical conclusions because they cannot build a sufficiently abstract model of certain problems that allows them to grasp their true nature. In other words, the whole class of questions of the kind of "how many r's in strawberry" or "do I take the car to the car wash?" would be answered correctly and reliably.

Re: Do LLMs pass the mirror test?

#48
> LLMs have seen humans act like conscious beings all over their training data because humans acting like conscious beings IS their training data.

How do we know that humans don't learn how to act conscious by observing other humans who act conscious?

Consciousness doesn't have a precise definition, but if you ask someone to describe it, there is a good chance that the description will include the concept of internal monologue.

The problem is that "internal" monologue is completely meaningless if you never heard an external monologue.

Also, people usually describe internal monologue as something that uses language and language is impossible to learn without communicating with other humans or at least observing other humans.

What I'm saying is that "well, LLM just pretends to be conscious, because it observed humans acting like conscious beings" doesn't really helps us to create a meaningful distinction between human consciousness and machine "consciousness", because same can be argued about us.

We don't know if feral children [0] are conscious and we don't know how to check it.

[0] https://en.wikipedia.org/wiki/Feral_child

Re: Do LLMs pass the mirror test?

#49
post #44

Earlier quoted context omitted.

If you don’t mind actually taking a few more words to be more specific that would be helpful because what you’re saying doesn’t really make sense at all. You don’t need to trust that the reasoning traces are all faithful representation of an internal reasoning trace. Plenty of other ways to probe models (see anthropics work using circuit tracing).

What else is there to say? LLMs can at most regurgitate approximations of human reasoning steps in the limited forms in which they may be expressed in the training data or interpolations thereof. That's the core essence of what they are. There is no proper reasoning to be found.

"at most" is wrong. RL with verifiable rewards takes you beyond quality and skills represented in training data, I'm not aware of meaningful fundamental limits here if you scale compute enough even though right now it's highly sample inefficient.

Since you refuse to actually define what you consider to be reasoning let me at least put one out there: a system exhibits reasoning when an answer depends on nontrivial intermediate computation over the problem. If you find problems with this, fine, but just make an effort to contribute an alternative.

If you increase test time compute you get better performance. If the model was just "interpolating" this wouldn't really work would it? Models can do FrontierMath expert problems (unpublished, expert authored, peer reviewed math problems) that require an insane amount of compositional reasoning. If they were regurgitating training data, that wouldn't really work would it? Chain of thought, while not always faithful to internal computation, improves performance. If the models were just regurgitating information, it wouldn't work that well would it?

"regurgitating training data" is also of course misleading. Yea they can memorize parts of the training data, but they generalize very well.

Re: Do LLMs pass the mirror test?

#50
post #47

Earlier quoted context omitted.

How do you define reasoning? What does a system have to functionally do in order to qualify for it?

Reasoning includes things like proper use of logic. LLMs have been repeatedly shown to fail horribly at this. They consistently fail at drawing basic logical conclusions because they cannot build a sufficiently abstract model of certain problems that allows them to grasp their true nature. In other words, the whole class of questions of the kind of "how many r's in strawberry" or "do I take the car to the car wash?"…

> Reasoning includes things like proper use of logic. LLMs have been repeatedly shown to fail horribly at this.

That models cannot do ALL logic problems does not mean that they cannot properly use logic...they can write Lean-verified theorems. How is that not logic?

> They consistently fail at drawing basic logical conclusions because they cannot build a sufficiently abstract model of certain problems that allows them to grasp their true nature.

What does their "grasp[ing] their true nature" have anything to do with what they can do?

> In other words, the whole class of questions of the kind of "how many r's in strawberry" or "do I take the car to the car wash?" would be answered correctly and reliably.

Again, just because you have interesting failure modes or brittleness does not mean they do not reason.

Post reply on HN