Live data from Hacker News

Mechanistic interpretability researchers applying causality theory to LLMs

cacm.acm.org

11–20 of 101 posts

Re: Mechanistic interpretability researchers applying causality theory to LLMs

#11

Earlier quoted context omitted.

Do they ?

Of course they do, how else do you think they manage to implement new features in large codebases, or to prove new theorems? But you don't even have to assume they do because of the results- you can read their chain of thought.

The Eliza effect.

Re: Mechanistic interpretability researchers applying causality theory to LLMs

#12
post #5

Earlier quoted context omitted.

The article answers this question, at least to the extent it can be answered, at this time. We see some signs of reasoning, but also we understand little about how they work.

Do we see actual signs of reasoning or is it anthropomorphism? We have an innate tendency to do so as humans.

Yes, we do see signs of actual reasoning, see the papers linked in the article. (There are many others too.)

Yes, we have a tendency to anthropomorphize, but (most) researchers are aware of this.

Re: Mechanistic interpretability researchers applying causality theory to LLMs

#13
post #61

[stub for offtopicness] [[All: please don't post shallow-generic reactions to baity titles. Those are basically the same thing, a la https://en.wikipedia.org/wiki/Rubin_vase , and we're trying for something more substantive here.]]

They don't reason.

Re: Mechanistic interpretability researchers applying causality theory to LLMs

#14
post #61

[stub for offtopicness] [[All: please don't post shallow-generic reactions to baity titles. Those are basically the same thing, a la https://en.wikipedia.org/wiki/Rubin_vase , and we're trying for something more substantive here.]]

My toaster doesn't reason, and neither do the current clankers.

Re: Mechanistic interpretability researchers applying causality theory to LLMs

#15

Earlier quoted context omitted.

Do they ?

Yes, there is an LLM feature that we have anthropomorphized as "reasoning" or "thinking", where an LLM has a scratch space where it can dump tokens that help to improve the final output.

> that help to improve the final output

Do they actually help? Are you sure?

Re: Mechanistic interpretability researchers applying causality theory to LLMs

#16

Earlier quoted context omitted.

Of course they do, how else do you think they manage to implement new features in large codebases, or to prove new theorems? But you don't even have to assume they do because of the results- you can read their chain of thought.

[flagged]

For the love of all that is sacred, please stop doing this. I'm begging you. The whole social media landscape is dying and you are creating a throwaway to participate in ruining this small corner. I assume this is not your first. And no one is convinced by this! The guidelines are there for your benefit as well. You achieve nothing but hastening the destruction of one of the last half-decent communities. Sorry for the melodrama.

Re: Mechanistic interpretability researchers applying causality theory to LLMs

#17
One plausible reason I thought of that we may not understand neural nets is that by their nature their power grows with ever-more complex connections and weights.

So it is like the opposite of logical systems, in that the very design of neural net architecture is a mess of parameter "spaghetti code" which renders the entire thing a metaphorical encrypted black box. The more powerful an AI/AGI the more this would be the case, and this is analogous a complexity curve.

And so any effort to make sense of such black box computation would be like trying to reverse entropy, analogous to trying to recover information lost in waste heat. And that could be one fundamental barrier to understanding both human and artificial brains alike, relative to their internal complexity.

(Just thinking aloud my handwavy pet theory recently, I am not an expert and could be totally mistaken on this)

Re: Mechanistic interpretability researchers applying causality theory to LLMs

#18
post #61

[stub for offtopicness] [[All: please don't post shallow-generic reactions to baity titles. Those are basically the same thing, a la https://en.wikipedia.org/wiki/Rubin_vase , and we're trying for something more substantive here.]]

They dont. They have input that runs through a invisible stochastic canyon. As long as there is previous experience the stochastic canyon never ends. If there is none or isignificant one, or it runs out of tokkens, it hallucinates and the illusion falls apart. There is no reasoning, just the invisible grand canyon of all of human experience and knowledge. PS: try to get it to retell you a clichee movie or book and you can see life near the end, how the delta of all the same movies opens up into wildly different endings.

To advance further it would need the ability to abstract away the general situation shape and pattern recognize similar situations.

Re: Mechanistic interpretability researchers applying causality theory to LLMs

#19
post #12

Earlier quoted context omitted.

Do we see actual signs of reasoning or is it anthropomorphism? We have an innate tendency to do so as humans.

Yes, we do see signs of actual reasoning, see the papers linked in the article. (There are many others too.) Yes, we have a tendency to anthropomorphize, but (most) researchers are aware of this.

The papers linked in the article discuss the mechanical operations that simulate reasoning. Intelligence is data efficiency and I don't see a strong argument that reasoning can exist if it requires a world's worth of data.

That doesn't mean that simulated reasoning isn't useful, it's wildly useful. But a thing is not its simulation.

Post reply on HN