Live data from Hacker News

Irrelevant facts about cats added to math problems increase LLM errors by 300%

science.org

131–140 of 270 posts

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#131

Earlier quoted context omitted.

Is the model thinking what is cat doing here? Then start thinking it is being tested?

I have no clue what the model is thinking, and as far as I can tell the paper also makes no attempt at answering that. It's also not really the point, the point is more that the claim in the paper that humans would be unaffected is unsubstantiated and highly suspect. I'd even say more likely wrong than right

> It's also not really the point, the point is more that the claim in the paper that humans would be unaffected is unsubstantiated and highly suspect.

I think the question that adds a random cat factoid at the end is going to trip up a lot fewer humans than you think. At the very least, they could attempt to tell you after the fact why they thought it was relevant.

And ignoring that, obviously we should be holding these LLMs to a higher standard than “human with extraordinary intelligence and encyclopedic knowledge that can get tripped up by a few irrelevant words in a prompt.” Like, that should _never_ happen if these tools are what they’re claimed to be.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#132
post #131

Earlier quoted context omitted.

I have no clue what the model is thinking, and as far as I can tell the paper also makes no attempt at answering that. It's also not really the point, the point is more that the claim in the paper that humans would be unaffected is unsubstantiated and highly suspect. I'd even say more likely wrong than right

> It's also not really the point, the point is more that the claim in the paper that humans would be unaffected is unsubstantiated and highly suspect. I think the question that adds a random cat factoid at the end is going to trip up a lot fewer humans than you think. At the very least, they could attempt to tell you after the fact why they thought it was relevant. And ignoring that, obviously we should be holding th…

I'm sure humans would be affected in some way. But not al all the same way an LLM would.

A human would probably note it as a trick in their reply.

The way LLMs work it could bias their replies in weird ways by changing their replies in unexpected ways beyond seeing it as a trick.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#133

Earlier quoted context omitted.

Ya, I specifically remember solving word problems in school / college and getting distracted by irrelevant details. Usually I would get distracted by stuff that _seemed_ like it should be used, so maybe cat facts would be fine for me to tease out, but in general I don't think I'm good at ignoring extraneous information. Edit: To be fair, in the example provided, the cat fact is _exceptionally_ extraneous, and even fl…

I had always assumed that the extraneous information was part of the test. You have to know/understand the concept well enough to know that the information was extraneous.

From what I remember of school, extraneous information was rarely included and the teachers who did add extraneous information seemed to do it maliciously.

There was one math class at a private school I attended that was the exception. The textbook had identifying relevant information as part of several chapters.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#134
post #80

There is more than one comment here asserting that the authors should have done a parallel comparison study against humans on the same question bank as if the study authors had set out to investigate whether humans or LLMs reason better in this situation. The authors do include the claim that humans would immediately disregard this information and maybe some would and some wouldn't that could be debated and seemingly…

to put it in better context, the problem is "does having a ton of MCP tool definitions available ruin the LLM's ability to design and write the correct code?"

and the answer seems to be yes. its a very actionable result about keeping tool details out of the context if they arent immediately useful

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#135
post #127

Earlier quoted context omitted.

The problem with your reasoning is that some humans cannot solve the problem even without the irrelevant info about cats. We can easily cherry pick our humans to fit any hypothesis about humans, because there are dumb humans. The issue is that AI models which, on the surface, appear to be similar to the smarter quantile of humans in solving certain problems, become confused in ways that humans in that problem-solving…

That's obviously because the brain is not generally intelligent it's just retrieving concepts from a high-dimensional statistically fit function. The extra info injects noise into the calculation which confounds it.

Yes, how... obvious?

I don't know, do we even know how the brain works? Like, definitively? Because I'm pretty sure we don't.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#136

Earlier quoted context omitted.

Yeah you're right, if that human is 5 years old or has crippling ADHD.

You think too highly of humans. Humans are not reliable. For every "no human would make this kind of mistake", you can find dozens to hundreds of thousands of instances of humans making this kind of mistake.

That's just because there's a lot of humans and we're doing a lot of things, all the time.

Humans are pretty good at not making mistakes in high-reasoning scenarios. The problem is that humans make mistakes in everything pretty constantly. Like, even saying a word - people say the wrong word all the time.

So when we look at really easy tasks that can be trivially automated, like say adding 2 + 2, we say "humans are so stupid! Computer is smart!".

Because humans get 2 + 2 wrong 1% of the time, but computers always get it right.

But, as we know, this isn't how it works. Actually, humans are much smarter than computers, and it's not even close. Because intelligence is multi-dimensional. The thing is, that failure rate for humans stays pretty constant as the complexity of the task increases, to a degree. Whereas computers start failing more and more, and quickly. It's a very, VERY sharp cliff for algorithms.

LLMs take the cliff further, but they do not eliminate it.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#137
post #3

> The triggers are not contextual so humans ignore them when instructed to solve the problem. Do they? I've found humans to be quite poor at ignoring irrelevant information, even when it isn't about cats. I would have insisted on a human control group to compare the results with.

Guilty. I remember taking an aptitude test in primary school, and choosing an answer based on my familiarity with the subject in the math test (IIRC the question mentioned the space shuttle) instead of actually attempting to solve the problem. I got cleanly filtered on that test.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#138

Earlier quoted context omitted.

As someone who has written and graded a lot of University exams, I'm sure a decent number of students would write the wrong answer to that. A bunch of students would write 5 (adding all the numbers). Others would write "3 apples and 2 cats", which is technically not what I'm looking for (but personally I would give full marks for, some wouldn't). Many students clear try to answer exams by pattern matching, and I've s…

When you try wing your way through a question by pattern matching, then you are not applying intelligence. Your interests lie elsewhere and so you are just fumbling your way through the activity at hand just to get through it.

This is something that the rise of LLMs has highlighted for me. Sometimes, we don't care to apply our intelligence to a problem. I've come to think of myself as "acting like an LLM" when I do this.

It reminds me of Kahneman's "system 1" (fast) and "system 2" (slow) thinking. LLMs are system 1 - fast, intuitive, instinctual. Humans often think that way. But we can also break out system 2 when we choose to, and apply logic, reason, etc.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#140
post #67

Earlier quoted context omitted.

> which really wouldn't confuse most humans And i think it would. I think a lot of people would ask the invigilator to see if something is wrong with the test, or maybe answer both questions, or write a short answer on the cat question too or get confused and give up. That is the kind of question where if it were put to a test I would expect kids to start squirming, looking at each other and the teacher, right as the…

Yeah you're right, if that human is 5 years old or has crippling ADHD.

You can argue until the cows come home. The point is that they claim without evidence that humans are not suspectible to this kind of distraction.

If they want to estabilish this as a fact there is a trivialy easy experiment they can conduct.

“Someone on hacker news strongly feels it is true, and is willing to argue the case with witty comments.” is not how scientific knowledge is estabilished. We either have done the experiments and have the data, or we don’t.

Post reply on HN