Live data from Hacker News

Mechanistic interpretability researchers applying causality theory to LLMs

cacm.acm.org

61–70 of 101 posts

Re: Mechanistic interpretability researchers applying causality theory to LLMs

#62
post #48

I personally would not look for the way they reason in the weights, at least not directly. In principle I could replace a large language model with a map from all possible input strings to output token or output token distribution without any weights. I have a hard time imagining how you would even tell, at the level of weights and activations, if the next token being the is the result of some proper reasoning or a h…

I also suspect the learned patterns are not necessarily efficient, though might be by accident.

One could “learn” addition by memorizing a truth table instead of understanding the concept… The truth table itself wouldn’t have much meaning.

Re: Mechanistic interpretability researchers applying causality theory to LLMs

#63
post #40

This article is not about "reasoning" in the abstract, philosophical sense but is talking about "mechanistic interpretability" research. The title is more like, "can we understand if the 'knowledge' encoded into a neural networks actually corresponds to reasoning-like concepts" and doing that with actual experiments like tweaking weights and activations. There's an interesting example where researchers saw a model ap…

Thanks - I've attempted to put that in the title above, in the hope of representing the article accurately.

(The trouble with a baity title like "Can We Understand How Large Language Models Reason?" is that it generates a barrage of shallow, reflexive responses having little to do with the article. What we want on HN are curious, reflexive responses instead - https://hn.algolia.com/?dateRange=all&page=0&prefix=true&sor....)

Re: Mechanistic interpretability researchers applying causality theory to LLMs

#64
post #8
post #61

[stub for offtopicness] [[All: please don't post shallow-generic reactions to baity titles. Those are basically the same thing, a la https://en.wikipedia.org/wiki/Rubin_vase , and we're trying for something more substantive here.]]

Clickbait article title. The article body does not presume they reason.

We've edited the title now in the hope of nudging the discussion in a more substantive direction.

Re: Mechanistic interpretability researchers applying causality theory to LLMs

#69
post #53

Earlier quoted context omitted.

The question of whether a SAT solver can reason is about as interesting as the question of whether a submarine can swim. (EWD867, EWD898)

I think you are missing the point of that statement It is a claim that swimming is a word that defines a context. It is an explicit statement that the question of whether a submarine can swim has nothing to do with the capability of the submarine. If you are asking which pigeon hole we are putting something into, the answer is "The one we put it into". This is what make the question uninteresting. If you are asking w…

The statement takes meaning-as-use as a given, sure, but I think the point of the statement is that people are arguing over an uninteresting question / taking meaningless positions about a meaningless issue, rather than "hey, words are moves in a language game!". I referenced two EWDs, which provide the original statements in context (though I can't find the widely-quoted wording anywhere: I thought I remembered it being in EWD1035, but apparently not). If you think my understanding of what Dijkstra meant was wrong, could you explain further, please?

Re: Mechanistic interpretability researchers applying causality theory to LLMs

#70
post #53

Earlier quoted context omitted.

I think you are missing the point of that statement It is a claim that swimming is a word that defines a context. It is an explicit statement that the question of whether a submarine can swim has nothing to do with the capability of the submarine. If you are asking which pigeon hole we are putting something into, the answer is "The one we put it into". This is what make the question uninteresting. If you are asking w…

The statement takes meaning-as-use as a given, sure, but I think the point of the statement is that people are arguing over an uninteresting question / taking meaningless positions about a meaningless issue, rather than "hey, words are moves in a language game!". I referenced two EWDs, which provide the original statements in context (though I can't find the widely-quoted wording anywhere: I thought I remembered it b…

I take your point that this seems to be the implication of what Dijkstra was getting at. But the term itself is not evocative for that reason. It resonates because people clearly don't think it is a meaningful question as to whether or not submarines swim because the term swim itself implies the categorisation I mentioned in my post above.

I do not know whether Dijkstra understood this distinction and was using it to disingenuously imply that the limitation was on the target and not the categorisation. He may have just felt it resonate with himself and failed to explore why.

Dijkstra immediately before using the term throws shade on serious thinkers engaging in a topic seriously. He personally seemed to want to dismiss the issue out of hand. As such I don't think there is any real value in his opinion on the matter. A recognition of how people did take it seriously and a considered rebuttal would be worthwhile. Declaring it uninteresting and failing to engage in the arguments is simply opting out of the debate.

Post reply on HN