Live data from Hacker News

Irrelevant facts about cats added to math problems increase LLM errors by 300%

science.org

211–220 of 270 posts

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#211

Oh no, just when we finally got them to properly count the number of "R"s in "strawberry"...

Hopefully these cases will get viral to the general public, so that everyone becomes more aware that despite the words "intelligence", "reasoning", "inference" being used and misused, in the end it is no more than a magic trick, an illusion of intelligence.

That being said, I also have hopes in that same technology for its "correlation engine" aspect. A few decades ago I read an article about expert systems; it mentioned that in the future, there would be specialists that would interview experts in order to "extract knowledge" and formalize it in first order logic for the expert system. I was in my late teens at that time, but I instantly thought it wasn't going to fly: way too expensive.

I think that LLMs can be the answer to that problem. One often reminds that "correlation is not causation", but it is nonetheless how we got there; it is the best heuristic we have.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#212
Why do we keep having these LLM studies that are completely unsurprising. Yes, the probabilistic text generator is more likely to output a correct answer when the input more closely matches its training sources than when you add random noise to the prompt. They don’t actually “understand” maths. It’s worrying how much research seems to operate from the premise that they do.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#213

Earlier quoted context omitted.

> models deployed in critical applications such as finance, law, and healthcare. We went really quickly from "obviously noone will ever use these models for important things" to "we will at the first opportunity, so please at least try to limit the damage by making the models better"...

Today someone who is routinely drug tested at work is being replaced by a hallucinating LLM.

To be fair, the AI probably hallucinates more efficiently than the human.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#214

Earlier quoted context omitted.

> We need to move past the humans vs ai discourse it's getting tired. You want a moratorium on comparing AI to other form of intelligence because you think it's tired? If I'm understanding you correctly, that's one of the worst takes on AI I think I've ever seen. The whole point of AI is to create an intelligence modeled on humans and to compare it to humans. Most people who talk about AI have no idea what the psycho…

>The whole point of AI is to create an intelligence modeled on humans and to compare it to humans. According to who? Everyone who's anyone is trying to create highly autonomous systems that do useful work. That's completely unrelated to modeling them on humans or comparing them to humans.

What do you imagine the purpose of these models' development is if not to rival or exceed human capabilities?

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#215

I'm going to write duck facts in my next online argument to stave off the LLMs. Ducks start laying when they’re 4-8 months old, or during their first spring.

As many as ten hundred thousand billion ducks are known to flock in semiannual migrations, but I think you'll find corpus distortion ineffective at any plausible scale. That egg has long since hatched.

> That egg has long since hatched.

I imagine there's entire companies in existence now, whose entire value proposition is clean human-generated data. At this point, the Internet as a data source is entirely and irrevokably polluted by large amounts of ducks and various other waterfowl from the Anseriformes order.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#216

This looks like it'll be useful for CAPTCHA purposes. According to the researchers, “the triggers are not contextual so humans ignore them when instructed to solve the problem”—but AIs do not. Not all humans, unfortunately: https://en.wikipedia.org/wiki/Age_of_the_captain

I tried the Age of the Captain on Gemini and ChatGPT and both game smarmy answers of "ahh this a classic gotcha". I managed to get ChatGPT to then do some interestng creative inference but Gemini decided to be boring.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#217
post #202

Earlier quoted context omitted.

> if they are going to be mass deployed in society This is the crucial point. The vision is massive scale usage of agents that have capabilities far beyond humans, but whose edge case behaviours are often more difficult to predict. "Humans would also get this wrong sometimes" is not compelling.

It's also off-the-charts implausible to say that our performance on adding up substantially degrades with the introduction of irrelevant information. Almost all cases of our use of arithmetic in daily life come with vast amounts of irrelevant information. Any person who looked at a restaurant table and couldn't review the bill because there were kid's drawings of cats on it would be severely mentally disabled, and ne…

> It's also off-the-charts implausible to say that our performance on adding up substantially degrades with the introduction of irrelevant information

Didn't you ever sit an exam next to a irresistibly gorgeous girl? Or haven't you ever gone to work in the middle of a personal crisis? Or filled out a form while people were rowing in your office? Or written code with a pneumatic drill and banging outdoors?

That's the kind of irrelevant information in our working context that will often degrade human performance. Can you really argue noise in a prompt is any different?

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#218
post #193

I try to be polite to the LLM and say e.g. thank you. Now I wonder if it is costing me quality.

Why be polite to a machine?

Because I want to be a polite person by default. It makes life nicer fot everyone involved and gives extra effect when I (rarely)choose not to be polite. I believe any interaction with anything is a little training, and I want to do it in the right direction.

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#219
post #69
post #41

Earlier quoted context omitted.

Read the article before commenting next time and you wont end up looking like a typical redditor.

“Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that". ” -- https://news.ycombinator.com/newsguidelines.html

Thanks, will stick to that in future

Re: Irrelevant facts about cats added to math problems increase LLM errors by 300%

#220

Earlier quoted context omitted.

It's also off-the-charts implausible to say that our performance on adding up substantially degrades with the introduction of irrelevant information. Almost all cases of our use of arithmetic in daily life come with vast amounts of irrelevant information. Any person who looked at a restaurant table and couldn't review the bill because there were kid's drawings of cats on it would be severely mentally disabled, and ne…

> It's also off-the-charts implausible to say that our performance on adding up substantially degrades with the introduction of irrelevant information Didn't you ever sit an exam next to a irresistibly gorgeous girl? Or haven't you ever gone to work in the middle of a personal crisis? Or filled out a form while people were rowing in your office? Or written code with a pneumatic drill and banging outdoors? That's the…

"Intelligence" is a metaphor used to describe LLMs (, AI) used by those who have never studied intelligence.

If you had studied intelligence as a science of systems which are intelligent (ie., animals, people, etc.) then this comparison would seem absurd to you; mendacious and designed to confound.

The desperation to find some scenario in which, at the most extreme superficial level, an intelligent agent "benchmarks like an LLM" is a pathology of thinking designed to lure the gullible into credulousness.

If an LLM is said to benchmark on arithmetic like a person doing math whilst being tortured, then the LLM cannot do math -- just as a person being tortured cannot. I cannot begin to think what this is supposed to show.

LLMs, and all statistical learners based on interpolating historical data, have a dramatic sensitivity to permuting their inputs such that they collapse in performance. A small permutation to the input is, if we must analogise, "like toturing a person to the point their mind ceases to function". Because these learners do not have representations of the underlying problem domain which are fit to the "natural, composable, general" structures of that domain ---- they are just fragmaents of text data put in a blender. You'll get performance only when that blender isnt being nudged.

The reason one needs to harm a person to a point they are profoundly disabled and cannot think, to get this kind of performance -- is that at this point, a person cannot be said to be using their mind at all.

This is why the analogy holds in a very superficial way: because LLMs do not analogise to functioning minds; they are not minds at all.

Post reply on HN