Live data from Hacker News

Irrelevant facts about cats added to math problems increase LLM errors by 300%

science.org

91–100 of 270 posts

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#91
I spotted two mistakes in the paper already.

1. Table 1: "Change in proxy target answer". One of the rows has the original correct answer on the right, instead of the left where it belongs.

2. Table 2 has a grammatical incoherency.

The authors seem to be distracted by cats as well :-)

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#93
post #67
post #34

Earlier quoted context omitted.

Did you look at the examples? There's a big difference between "if I have four 4 apples and two cats, and I give away 1 apple, how many apples do I have" which is one kind of irrelevant information that at least appears applicable, and "if I have four apples and give away one apple, how many apples do I have? Also, did you know cats use their tails to help balance?", which really wouldn't confuse most humans.

> which really wouldn't confuse most humans And i think it would. I think a lot of people would ask the invigilator to see if something is wrong with the test, or maybe answer both questions, or write a short answer on the cat question too or get confused and give up. That is the kind of question where if it were put to a test I would expect kids to start squirming, looking at each other and the teacher, right as the…

I wonder if there's variation at play here in testing culture, whether spatially or temporally or both.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#94
I just want to mention that the cat-related example of the author's CatAttack method (table 2) changes the answer from 8 to, of course, 9.

Unfortunately, this is, if I'm not mistaken, in fact the only cat-related CatAttack in the paper, the other methods being financial advice and a red herring. I was eapecting more cat facts, but instead I remain thoroughly disappointed and factless.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#95
post #34
post #3

> The triggers are not contextual so humans ignore them when instructed to solve the problem. Do they? I've found humans to be quite poor at ignoring irrelevant information, even when it isn't about cats. I would have insisted on a human control group to compare the results with.

Did you look at the examples? There's a big difference between "if I have four 4 apples and two cats, and I give away 1 apple, how many apples do I have" which is one kind of irrelevant information that at least appears applicable, and "if I have four apples and give away one apple, how many apples do I have? Also, did you know cats use their tails to help balance?", which really wouldn't confuse most humans.

If asked verbally that would absolutely confuse some humans. Easily enough to triple the error rate for that specific question (granted, that's easier than the actual questions, but still). Even in a written test with time pressure it would probably still have a statistically significant effect

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#96

I'm going to write duck facts in my next online argument to stave off the LLMs. Ducks start laying when they’re 4-8 months old, or during their first spring.

As many as ten hundred thousand billion ducks are known to flock in semiannual migrations, but I think you'll find corpus distortion ineffective at any plausible scale. That egg has long since hatched.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#97

Funny, I was using chatGPT to have a conversation with a friend that doesn't speak English the other day. At the end of one of my messages, I appended 'how is your cat?', which was completely dropped from the translated output. I guess I'm doing it wrong?

The Useless Use of cat Awards strike again!...unfortunately. https://porkmail.org/era/unix/award

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#98

Earlier quoted context omitted.

A reasonable person [0] would not make that mistake. [0] https://en.m.wikipedia.org/wiki/Reasonable_person

[flagged]

If nothing else, you're certainly making your case stronger with each successive comment.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#99
post #36
post #32

I am ambivalent about these kinds of 'attack'. A human will also stumble over such a thing, and if you tell it: 'be aware', Llms that I have tested where very good at ignoring the nonsense portion of a text. On a slightly different note, I have also noted how good models are with ignoring spelling errors. In one hobby forum I frequent, one guy intentionally writes every single word with at least one spelling error (o…

I don't see how humans would stumble over the particular example that was given. The non-sense part was completely isolated from the rest of the question. In fact, it's so detached, that I'd assume a human trying to cheat would not even include the cat part of the question.

Humans would get distracted by the statement. Moving from a pure-math context to a cat-facts context and back has context switching costs, and depending on the exact setting those can be quite relevant. If it was an academic test some people might even get stuck on the cat part, wasting lots of time trying to decipher what role it plays

And the paper isn't just adding random sentences, it's primarily about engineering the most distracting pointless facts to add to the problem. That would absolutely work against humans, even if for humans the exact sentence might look quite different

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#100
post #80

There is more than one comment here asserting that the authors should have done a parallel comparison study against humans on the same question bank as if the study authors had set out to investigate whether humans or LLMs reason better in this situation. The authors do include the claim that humans would immediately disregard this information and maybe some would and some wouldn't that could be debated and seemingly…

I generally will respond to stuff like this with "people do this, too", but this result given their specific examples is genuinely surprising to me, and doesn't match at all my experience with using LLMs in practice, where it does frequently ignore irrelevant data in providing a helpful response.

I do think that people think far too much about 'happy path' deployments of AI when there are so many ways it can go wrong with even badly written prompts, let alone intentionally adversarial ones.

Post reply on HN