Live data from Hacker News

Simple tasks showing reasoning breakdown in state-of-the-art LLMs

arxiv.org

211–220 of 393 posts

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#211

Earlier quoted context omitted.

Given many people don’t have an inner monologue and function just fine, it’s more likely inner monologue is a product of the reasoning process and not it’s mechanism.

I think you’re using “inner monologue” too literally. It could be a progression of pictures, emotions, etc.

To make any progress on this question at all, we need first to come up with some definition of internal monologue. Even if we may need to modify it later, there has to be a starting point.

Otherwise, nothing can be established at all, because for any statement there always will be someone's understanding of "internal monologue" for which the statement is true, and someone's else understanding for which the statement is false...

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#212

Earlier quoted context omitted.

Humans train on continuous video . Even our most expensive models are, in terms of training set size, far behind what an infant processes in the first year of their life. EDIT: and it takes human children a couple years to reliably identify a cat. My 2.5 y.o. daughter still confuses cats with small dogs, despite living under one roof with a cat.

I contend that you could show any child old enough to communicate in basic English a photograph (so not live continuous video) of some obscure animal they've never seen before (say an Okapi) and they'd be able to easily identify another Okapi when seeing one at a zoo.

So you're just going to ignore the 5 years of continuous training? I'm not sure what point you're trying to make.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#213

Many of the datasets for the "benchmarks" on which the major public LLMs are assessed are clearly present in their training data, making them basically useless for establishing reliability of the models. Its fairly obvious that at least some of the improved scores from later generations of models are that this benchmark data is increasingly represented in the training data. A better way of assessing LLMs is waiting a…

MMLU is not a reasoning benchmark. It's a measure of how distributed and representative their training data was and how well it's able to recall (for lack of a better word) based on training epochs.

GPQA etc. test reasoning in some form, and you see the drastic change in score between the two for every model.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#214
> The breakdown is dramatic, as models also express strong overconfidence in their wrong solutions, while providing often non-sensical "reasoning"-like explanations akin to confabulations to justify and backup the validity of their clearly failed responses, making them sound plausible.

I like their use of confabulations instead of hallucinations. I think confabulate describes what LLMs are doing much better than hallucinate.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#215
post #87

Earlier quoted context omitted.

I appear to be reasoning at times but I have mostly no idea what I am talking about. I hit a bunch of words and concepts in the given context and thus kind of hallucinate sense. Given a few months of peace of mind and enough money for good enough food, I could actually learn to reason without sounding like a confused babelarian. Reasoning is mostly a human convention supported by human context that would have been a…

Yeah, I think these chatbots are just too sure of themselves. They only really do "system 1 thinking" and only do "system 2 thinking" if you prompt them to. If I ask gpt-4o the riddle in this paper and tell it to assume its reasoning contains possible logical inconsistencies and to come up with reasons why that might be then it does correctly identify the problems with its initial answer and arrives at the correct on…

If you had a prompt that reliably made the model perform better at all tasks, that would be useful. But if you have to manually tweak your prompts for every problem, and then manually verify that the answer is correct, that's not so useful.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#216
post #161

Earlier quoted context omitted.

With that definition even bacteria have inner monologue.

Can bacteria imagine pictures? Do they have emotions? Why does this matter? Stop being so pedantic. We're talking about a progression of ideas . Talking in your head is one form of ideas, but people can easily solve problems by imagining them.

Hmm, looks to me like just trading some words for others. Do bacteria have ideas? Does the navigating system in your car? How do you know?

We need to be at least somewhat pedantic, otherwise it's impossible to know what we are even talking about, and no way to establish anything.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#217
post #165

Earlier quoted context omitted.

Maybe, but symbolic thought can get pretty far away from what we generally call "language." I bet you can reason 1+3x=22 pretty easily without any words whatsoever, or the sound of one ascending octave after another, or the approximate G-force induced on your body if you take the next turn without applying the brakes. All of these forms of reasoning are true and useful calculations: when we talk about "intuition" wha…

Regarding “1+3x=22”, I’m actually not sure, the number words certainly appear in my head when solving the equation. But even then, I would count “1+3x=22” as constituting language. Perception of sound, G-forces, and dancing don’t perform type-2 reasoning by themselves, so I don’t think your argument applies there. Regarding your edit, no, I think the key aspect of the kind of reasoning we are missing in current AI is…

It is very difficult to have a discussion using words to discuss the semantics of non-word or non-symbolic semantics. I was pointing at several different plausible spaces for semiotics and how these spaces could be spaces for reasoning in the hopes that one of them might be relatable.

If you use words in your mind when you use math, and you use words in your mind when you make or listen to music, etc., then it is very difficult to find a common ground where it is possible to see that these other realms of thought are capable of not only prediction, but also producing evidence that leads to judgement. That is to say, the key aspects of "reasoning." I picked them because I thought they had broad enough appeal to be relatable, and because I do not personally hear words in my head when doing any of these activities, whether it's calculus or tango, but I still find calculus and tango to be places where reasoning occurs.

Some of them, like math or music, are closer to the kind of symbolic thought we use when we discuss things with words. Others, like the experience of g-forces, are not. I present them as a sliding scale between "word based" reasoning and "non-linguistic" reasoning. Perhaps you can think of a realm that better fits for your personal experience of intuition, and inspect whether these intuitions are capable of "real" reasoning in the absence of language, or whether intuition should never be trusted even when you have a great deal of experience in that area. Perhaps in your estimation, anything that cannot produce evidence that is articulable in word form is suspect.

Personally, I find all these methods, including language, to be suspect. I don't find language to be especially better at producing the kind of evidence for prediction, correct judgement, or discourse for reasoning than other methods, unless you reduce "reasoning" to tautologically require it.

One of the best tools of language is that we have writing that allows easy inspection or iteration of the written content; but these things are possible in other realms, too, it's just that we didn't have great tools for introspecting and iterating on their "ideas" except within our own minds. These days, those tools are readily available in many more realms of human insight.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#218
post #209

Earlier quoted context omitted.

Its not an AI hype. A hype is defined as something which gets oversold: "promote or publicize (a product or idea) intensively, often exaggerating its benefits." Just yesterday I visited a google cloud summit and one person from bosch told the audiance how they are now able to work with less external agencies like texting, graphicsdesigner and photographers for their materials. It already saves money, has real impacts…

Your post is complete hype, all about people saying things instead of showing things that've actually been done. For me, 2024 was the LLM exposed as basically pure hype year. There is no expert of any field I follow online where they're posting up results from AI tooling for any other reason than to show how awful it is. I consider myself an expert in software, and LLMs specifically have only caused me great pain. Ev…

Eat something and take a nap, you sound unhinged.

ChatGPT has nearly doubled my work output, most of my job is system admin infra type stuff and it's ridiculously good at troubleshooting odd issues.

Hopefully you can find a use case for it someday, until then, the rest of us will continue to be more productive.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#219
post #18

For anyone considering reading the paper and like me don't normally read papers like this, open the PDF and think you don't have time to read it due to its length. The main part of the paper is the first 10 pages and a fairly quick read. On to the topic here. This is an interesting example that they are using. It is fairly simplistic to understand as a human (even if we may be inclined to quickly jump to the wrong co…

The problem is a good chunk of the global population is also not reasoning and thinking in any sense of the word. Logical reasoning is a higher order skill that often requires formal training. It's not a natural ability for human beings.

In "any sense" of the word? Surely anyone who adjusts their behavior when they get undesired or unexpected results is reasoning and thinking. And since most activities are mediated by thought of some kind, most people are reasoning and thinking otherwise they would never recover from even simple mistakes, like walking east when they need to go north.

Saying they're "not thinking in any sense of the word" because they can't solve predicate logic problems from a college textbook is a rather odd claim. Surely those things arise from reasoning and thinking, rather than the other way around.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#220
post #18

For anyone considering reading the paper and like me don't normally read papers like this, open the PDF and think you don't have time to read it due to its length. The main part of the paper is the first 10 pages and a fairly quick read. On to the topic here. This is an interesting example that they are using. It is fairly simplistic to understand as a human (even if we may be inclined to quickly jump to the wrong co…

Its not an AI hype. A hype is defined as something which gets oversold: "promote or publicize (a product or idea) intensively, often exaggerating its benefits." Just yesterday I visited a google cloud summit and one person from bosch told the audiance how they are now able to work with less external agencies like texting, graphicsdesigner and photographers for their materials. It already saves money, has real impacts…

There is no feedback. You cannot create new knowledge out of thin air.
Post reply on HN