Live data from Hacker News

Is the reversal curse in LLMs real?

andrewmayne.com

111–120 of 211 posts

Re: Is the reversal curse in LLMs real?

#111

I played along the example of the chancellor question. If you deviate a little from the examples in the article GPT-4 gets it all wrong: Who is the eighth Federal Chancellor of the Federal Republic of Germany? - Olaf Scholz (wrong, Angela Merkel) https://chat.openai.com/share/937795ea-bd91-43ee-bd76-1a125f... Who was the eighth Federal Chancellor of the Federal Republic of Germany? - Helmut Kohl (wrong, Angela Merkel…

>Who is the eighth Federal Chancellor of the Federal Republic of Germany? - Olaf Scholz (wrong, Angela Merkel)

Technically there's no right answer to this question. "Is" implies present. But Angela Merkel isn't the present chancellor. Olaf Scholz is. But he's not the eigth.

Re: Is the reversal curse in LLMs real?

#112

Earlier quoted context omitted.

smart is in quotes here for a reason. It's 'dumb' in relation to many people's expectations. >Additionally, the training loop (which you call 'dumb') is just...a teacher-forced version of inference, which you call 'smart'? What powers In context learning is not very well understood but it doesn't appear to be or really work exactly like just a non teacher-forced version of training. There are qualitative differences.…

> What powers In context learning is not very well understood but it doesn't appear to be or really work exactly like just a non teacher-forced version of training. I mean, yes. One distills information from a training set into a compressed representation, and the other generates a compressed representation that yields (more or less) fixed state space attractors. It's just inducing a bias over the state space of the…

I guess if we'd automatically feed every such conversation back to the training set, it would be proper "in-context learning", and it would work very much like humans learn - inferring new facts from recalled facts, and memorizing the inferences that come up repeatedly or otherwise feel important.

Re: Is the reversal curse in LLMs real?

#113
post #8

I fall somewhere on the skeptic side of the LLM spectrum. But this "flaw" just does not seem to have the force that its proponents seem to think it does, unless I'm missing something significant. Simply because in the context of natural language (Rather than formal logical statements), "A is B" does not imply "B is A" in the first place. "Is" can encompass a wide variety of logical relationships in colloquial usage,…

This is explained very clearly in the paper. They are not looking at all sentences of the form "A is B":

>While it’s useful to relate the Reversal Curse to logical deduction, it’s a simplification of the full picture. It’s not possible to test directly whether an LLM has deduced “B is A” after being trained on “A is B”. LLMs are trained to predict what humans would write and not what is true (Lin et al., 2022). So even if an LLM had inferred “B is A”, it might not “tell us” when prompted. Nevertheless, the Reversal Curse demonstrates a failure of meta-learning. Sentences of the form “ is ” and “ is ” often co-occur in pretraining datasets; if the former appears in a dataset, the latter is more likely to appear.4 This is because humans often vary the order of elements in a sentence or paragraph.5 Thus, a good meta-learner would increase the probability of an instance of “ is ” after being trained on “ is ” . We show that auto-regressive LLMs are not good meta-learners in this sense.

Re: Is the reversal curse in LLMs real?

#114

I played along the example of the chancellor question. If you deviate a little from the examples in the article GPT-4 gets it all wrong: Who is the eighth Federal Chancellor of the Federal Republic of Germany? - Olaf Scholz (wrong, Angela Merkel) https://chat.openai.com/share/937795ea-bd91-43ee-bd76-1a125f... Who was the eighth Federal Chancellor of the Federal Republic of Germany? - Helmut Kohl (wrong, Angela Merkel…

as others pointed out, it's important how you ask and what you think the truth is. https://chat.openai.com/share/cc688592-236a-499b-82e1-81bdaa...

Re: Is the reversal curse in LLMs real?

#115

Earlier quoted context omitted.

Nah. The context of all of this is people claiming LLMs are AI. If so they would be able to reverse A is B, and also know when not to, so your example is irrelevant. The paper shows that it cannot reverse. Then the blog posts goes around different disingenuous ways of making excuses for why it can't reverse at the same time as trying to show that it can reverse in the Tom Cruise training example (extremely disingenuo…

I still call bullshit on this. Humans don't do reversals for free either. We may memorize the reversal, or first, infer it . That's an extra step, often very cognitively challenging. I'd like to see experiments showing how "reversal curse" fares when LLM is allowed to make the inference. Like: "Please complete the sentence: [your B->A] - but reason through the process first. Start with identifying the subject, then w…

This simply isn't true of humans for the kinds of examples used in the paper. If I read "Olaf Scholz was the ninth Chancellor of Germany" then I have no trouble reversing that and determining that the 9th chancellor of Germany was Olaf Scholz. This is not an inference that's 'cognitively challenging'.

Humans may sometimes fail to make simple inferences of this sort when building their factual databases, but they don't systematically fail to do so.

Also, you may have missed this part of the paper:

>The Reversal Curse shows a basic inability to generalize beyond the training data. Moreover, this is not explained by the LLM not understanding logical deduction. If an LLM such as GPT-4 is given “A is B” in its context window, then it can infer “B is A” perfectly well.

The paper does not claim that GPT-4 cannot perform logical deduction, but only that it does not appear to make use of it when generalizing its training data.

Re: Is the reversal curse in LLMs real?

#116
This is a great article with more tips on how to prompt and fine tune LLMs than most articles that focus on this topic. Also has surprising insights which are food for thought. Like: Would it help to attach facts to well-learned „nodes“ (like Tom Cruise in his example) when fine tuning? In the sense of repurposing well-known nodes. Looking forward to reading more from Andrew.

Re: Is the reversal curse in LLMs real?

#117

Earlier quoted context omitted.

I don't think this is the right explanation, or really that relevant, the suggestions from 'og_kalu and in the article sound more accurate to me. It seems like understanding when "is" is reversible is pretty core to the capabilities of the model, but that's different than having a lot of facts memorized. For instance, a model should be able to answer "who is the star of Mission Impossible" with "Tom Cruise" based on…

I mean transformers are definitely sequentially biased. There's also human speech bias. But I think it's pretty clear that humans are generally invariant to this kind of prompting as well (given that they have the knowledge. More in a different comment). My surprisal is far higher that the reversal "curse" is considered controversial or even surprising than it was when that sensational tweet dropped. It feels pretty…

> But I think it's pretty clear that humans are generally invariant to this kind of prompting as well (given that they have the knowledge. More in a different comment).

But are they? Unless I'm misunderstanding what you mean here, I'd say the opposite is clearly the case - learning "A is B" doesn't automatically mean learning the reversal; you either have to learn it from external source, or infer and memorize it - both involve an extra cognitive effort. This maps to LLMs as training, or "in-context learning" and then adding the inference to the training set.

Re: Is the reversal curse in LLMs real?

#118
Luckily this curse has no significant practical impact on accuracy.

Who knows who Mary Lee Pfeiffer is, but does not already know her son is Tom Cruise?

And if such a person exist, do you want to give them a correct answer, or talk about how Mary Lee Pfeiffer is not a notable person with a Wikipedia page?

Re: Is the reversal curse in LLMs real?

#119
post #115

Earlier quoted context omitted.

I still call bullshit on this. Humans don't do reversals for free either. We may memorize the reversal, or first, infer it . That's an extra step, often very cognitively challenging. I'd like to see experiments showing how "reversal curse" fares when LLM is allowed to make the inference. Like: "Please complete the sentence: [your B->A] - but reason through the process first. Start with identifying the subject, then w…

This simply isn't true of humans for the kinds of examples used in the paper. If I read "Olaf Scholz was the ninth Chancellor of Germany" then I have no trouble reversing that and determining that the 9th chancellor of Germany was Olaf Scholz. This is not an inference that's 'cognitively challenging'. Humans may sometimes fail to make simple inferences of this sort when building their factual databases, but they don'…

> This simply isn't true of humans for the kinds of examples used in the paper. If I read "Olaf Scholz was the ninth Chancellor of Germany" then I have no trouble reversing that and determining that the 9th chancellor of Germany was Olaf Scholz. This is not an inference that's 'cognitively challenging'.

That's in-context though. LLMs don't fail in-context either.

For a better comparison, recall your school experience, say with history lessons, or geography lessons - where you would cram a hundred "A is B" relationships, and then take a test that demanded you know the reversals. Not as easy.

The "reversal curse" failures I've seen with LLMs are very similar to asking a random person, out of a blue, some unusual reversal of some random fact they ought to know, and then being surprised they can't answer quickly.

> The paper does not claim that GPT-4 cannot perform logical deduction, but only that it does not appear to make use of it when generalizing its training data.

Well, neither can humans when cramming, if you don't give them time to pause and think about what they're learning. I believe the equivalent is happening here - LLMs can perform logical deductions, but at no point in the training process is this capability used.

Re: Is the reversal curse in LLMs real?

#120
I don't agree with authors claim.

There is a plant which mimics leaves of nearby plants discussed here https://news.ycombinator.com/item?id=31301454 If you ask GPT-4 what this plant is known for, it will tell you correctly. But if you ask in any number of ways to tell the name of plant which mimic leaves, it will always give incorrect answer.

Post reply on HN