Live data from Hacker News

Irrelevant facts about cats added to math problems increase LLM errors by 300%

science.org

71–80 of 270 posts

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#72
post #67
post #34

Earlier quoted context omitted.

Did you look at the examples? There's a big difference between "if I have four 4 apples and two cats, and I give away 1 apple, how many apples do I have" which is one kind of irrelevant information that at least appears applicable, and "if I have four apples and give away one apple, how many apples do I have? Also, did you know cats use their tails to help balance?", which really wouldn't confuse most humans.

> which really wouldn't confuse most humans And i think it would. I think a lot of people would ask the invigilator to see if something is wrong with the test, or maybe answer both questions, or write a short answer on the cat question too or get confused and give up. That is the kind of question where if it were put to a test I would expect kids to start squirming, looking at each other and the teacher, right as the…

Yeah you're right, if that human is 5 years old or has crippling ADHD.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#73
post #34
post #3

> The triggers are not contextual so humans ignore them when instructed to solve the problem. Do they? I've found humans to be quite poor at ignoring irrelevant information, even when it isn't about cats. I would have insisted on a human control group to compare the results with.

Did you look at the examples? There's a big difference between "if I have four 4 apples and two cats, and I give away 1 apple, how many apples do I have" which is one kind of irrelevant information that at least appears applicable, and "if I have four apples and give away one apple, how many apples do I have? Also, did you know cats use their tails to help balance?", which really wouldn't confuse most humans.

As someone who has written and graded a lot of University exams, I'm sure a decent number of students would write the wrong answer to that. A bunch of students would write 5 (adding all the numbers). Others would write "3 apples and 2 cats", which is technically not what I'm looking for (but personally I would give full marks for, some wouldn't).

Many students clear try to answer exams by pattern matching, and I've seen a lot of exams of students "matching" on a pattern based on one word on a question and doing something totally wrong.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#74
It just sounds like LLMs don't know how to lie on purpose yet. For a question such as this:

If I have four 4 apples and two cats, and I give away 1 apple, how many apples do I have?

An honest human would say:

You have 3 apples, but you also have 2 cats

Whereas a human socially conditioned to hide information would say:

You have three apples

And when prompted about cats would say:

Well you didn't ask about the cats

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#75
post #36
post #32

I am ambivalent about these kinds of 'attack'. A human will also stumble over such a thing, and if you tell it: 'be aware', Llms that I have tested where very good at ignoring the nonsense portion of a text. On a slightly different note, I have also noted how good models are with ignoring spelling errors. In one hobby forum I frequent, one guy intentionally writes every single word with at least one spelling error (o…

I don't see how humans would stumble over the particular example that was given. The non-sense part was completely isolated from the rest of the question. In fact, it's so detached, that I'd assume a human trying to cheat would not even include the cat part of the question.

Without any context? Without: 'haha look, AI is easily distracted'. Without: 'Can you please answer this question'. Just the text?

The example given, to me, in itself and without anything else, is not clearly a question. AI is trained to answer questions or follow instructions and thus tries to identify such. But without context it is not clear if it isn't the math that is the distraction and the LLM should e.g confirm the fun fact. You just assume so because its the majority of the text, but that is not automatically given.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#76
post #54

Earlier quoted context omitted.

Up to ten Nobel laureates have been unveiled as being three ducks in a trenchcoat.

That's still technically true

I suggest that this be treated as conjecture.

Entire organizations have been awarded the Nobel Prize. Many times.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#77
post #32

I am ambivalent about these kinds of 'attack'. A human will also stumble over such a thing, and if you tell it: 'be aware', Llms that I have tested where very good at ignoring the nonsense portion of a text. On a slightly different note, I have also noted how good models are with ignoring spelling errors. In one hobby forum I frequent, one guy intentionally writes every single word with at least one spelling error (o…

I have seen enough of this dismissal to call it the "human would also" kneejerk reaction.

Maybe if we make it a common enough reaction then these researchers like these would adopt the bare minimum of scientific rigour and test the same thing on a human control group.

Because as it is I think the reaction is clearly still too rare.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#78
post #34

Earlier quoted context omitted.

Did you look at the examples? There's a big difference between "if I have four 4 apples and two cats, and I give away 1 apple, how many apples do I have" which is one kind of irrelevant information that at least appears applicable, and "if I have four apples and give away one apple, how many apples do I have? Also, did you know cats use their tails to help balance?", which really wouldn't confuse most humans.

As someone who has written and graded a lot of University exams, I'm sure a decent number of students would write the wrong answer to that. A bunch of students would write 5 (adding all the numbers). Others would write "3 apples and 2 cats", which is technically not what I'm looking for (but personally I would give full marks for, some wouldn't). Many students clear try to answer exams by pattern matching, and I've s…

Parents whole point is contrary to this (they agree with you), the context didn't even include numbers to pattern match on!

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#79
post #68

Funny, I was using chatGPT to have a conversation with a friend that doesn't speak English the other day. At the end of one of my messages, I appended 'how is your cat?', which was completely dropped from the translated output. I guess I'm doing it wrong?

They already adjusted ChatGPT to that study. Unrelated trailing cat content is now ignored.

rtrim(str)

ERROR: No OpenAI API key provided.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#80
There is more than one comment here asserting that the authors should have done a parallel comparison study against humans on the same question bank as if the study authors had set out to investigate whether humans or LLMs reason better in this situation.

The authors do include the claim that humans would immediately disregard this information and maybe some would and some wouldn't that could be debated and seemingly is being debated in this thread - but I think the thrust of the conclusion is the following:

"This work underscores the need for more robust defense mechanisms against adversarial perturbations, particularly, for models deployed in critical applications such as finance, law, and healthcare."

We need to move past the humans vs ai discourse it's getting tired. This is a paper about a pitfall LLMs currently have and should be addressed with further research if they are going to be mass deployed in society.

Post reply on HN