Live data from Hacker News

Simple tasks showing reasoning breakdown in state-of-the-art LLMs

arxiv.org

301–310 of 393 posts

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#301

Earlier quoted context omitted.

I got it in 5 seconds, am I the singularity ?

We all are, but only in meat-space. We continue to improve ourselves much faster than evolution ever could. But what we are talking about here is the singularity in tech-space.

I don’t see such of a distinction between technology and us. We build, drive and continue our overflow this tech. It’s an extension of us inspired by how our own brains work.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#302

Earlier quoted context omitted.

It is how it works if you are replying to someone who claims "If you really think about what an LLM is you would think there is no way that leads to general purpose AI". The counter example is human beings are considered general purpose intelligence and we are complex, but fundamentally predictable systems (not by us today), with (as far as we can tell) deterministic outputs based on the state of the universe (includ…

Responding to an unsubstantiated claim with an unsubstantiated claim just adds another unsubstantiated claim. So far as I know, whether the universe behaves deterministically remains an unsolved question. Given that, your statement here would already be one of belief rather than fact, even before we get to the parentheticals. There is information here, but not about whether LLMs can develop into AGI.

Fine, you can ignore my previous comment, that's just my answer to the question that this discussion ultimately takes you to. But I feel like you are just sitting on the sidelines making strawmen and playing pedantic games instead of saying anything constructive.

The original comment said:

> If you really think about what an LLM is you would think there is no way that leads to general purpose AI.

This is an inflammatory way to state an extreme position on a well-discussed debate over whether next-token prediction can lead to general intelligence. The original commenter clearly believes it can't get you there. If you want to say that with any authority, you need to have an answer for what is different between what we consider general intelligence (for most people, this is simply human intelligence) and what models are capable of. This is the question at the heart of artificial intelligence.

I challenged them to explain their answer. I made no claims, I asked no one to prove anything wrong. If it is obvious that LLMs can't be AGI, the answer to how an LLM differs from human intelligence is also obvious, right?

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#303
post #18

For anyone considering reading the paper and like me don't normally read papers like this, open the PDF and think you don't have time to read it due to its length. The main part of the paper is the first 10 pages and a fairly quick read. On to the topic here. This is an interesting example that they are using. It is fairly simplistic to understand as a human (even if we may be inclined to quickly jump to the wrong co…

The problem is a good chunk of the global population is also not reasoning and thinking in any sense of the word. Logical reasoning is a higher order skill that often requires formal training. It's not a natural ability for human beings.

> Logical reasoning is a higher order skill that often requires formal training. It's not a natural ability for human beings.

I've read your comments here, and while I understand your point I think you have it backwards. The only reason we formed societies is because we evolved an innate a theory of mind to reason about how others might be thinking and feeling. That's reasoning. We have a natural ability to do limited arithmetic, otherwise we wouldn't be able to hunt, gather or track enough to endure winters, or keep track of our sheep or children for that matter. That's reasoning.

Reasoning is a natural ability for human beings, but we also carry a lot of evolutionary impulses that add a lot of noise to the decision process, eg. observation->judgment->[set of possible decisions], judgment has "reason" as one path that adds to the set of possible decisions, but there remain other paths we inherited from our evolutionary roots. Education is training that suppresses poorly calibrated judgment paths that lead to frequent mistakes in the decision set, but reasoning remains, thus education improves the signal to noise ratio of our decision making.

So I 100% disagree that an individual cannot separate cause and effect without training. They will just be worse at it than someone who is trained to filter out those impulses that lead us to jump to conclusions, eg. they will produce more noise / a larger set of possibilities than reason would allow.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#305

Earlier quoted context omitted.

The best compression is some form of understanding

That's a fascinating insight and it sound so true! Can you compress for me Van Gogh's Starry Night, please? I'd like to send a copy to my dear old mother who has never seen it. Please make sure when she decompresses the picture she misses none of the exquisite detail in that famous painting.

Okay yes so not really having an artists vocabulary I couldn't compress it as well as someone who has a better understanding of Starry Night. An artist that understands what makes Starry Night great could create a work that evokes similar feelings and emotions. I know this because Van Gogh created many similar works playing with the same techniques, colors, and subjects such as Cypresses in Starry Night and Starry Night over the Rhone. He was clearly working from a concise set of ideas and techniques which I would argue is understanding/compression.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#306

Earlier quoted context omitted.

In this particular case, is there any reason why we simply can't take their word for it? This is not a case of where if I say "weak" or "strong", most people pick strong because no one wants to be weak, even if the context is unknown (nuclear force for example).

Because we can't be sure whether two people interpret what "inner monologue" means and whether they think it describes a phenomenon that actually isn't different between them and other people. For example, I can think of interpretations of "I picture objects that I'm thinking about" that range from me not experiencing the phenomenon to me indeed experiencing the phenomenon. To say that you're not experiencing somethi…

And here I thought this was solved decades ago - I need to find the source, but I read about an old study where people describe their experience, and the answers were all over the "range from me not experiencing the phenomenon to me indeed experiencing the phenomenon".

Then again, it's trivially reproducible - people self-report all variants of inner monologue, including lack of it, whenever a question about it pops up on-line. Same is the case with imagination - aphantasia is a thing (I would know, I have it).

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#307
post #260

Citation 40 is the longest list of authors I have ever seen. That is one way to help all your friends get tenure.

https://arxiv.org/pdf/2206.04615

I guess this is the reason:

>> BIG-bench currently consists of 204 tasks, contributed by 450 authors across 132 institutions.

Regardless, I'm not citing a paper with a whole page of author names unless I'm allowed to shorten it significantly in the bibliography section (e.g. "Srivastava and 450 others").

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#308

As a non-coder I can get away with asking this: Why is it so hard to simulate reason? Logic and reason are based on rules. Then you add values to steer the conclusions based on the available data. Why not have separate systems for values and logic and memory working together as an AI brain to generate truly reasoned responses? You could even have adversarial parts that duke it out (left-wing vs right-wing, Jefferson…

>> As a non-coder I can get away with asking this: Why is it so hard to simulate reason?

It isn't. We know how to do reasoning with computers. The discussion about reasoning in LLMs is carried out in an echo chamber that ignores the prior work on reasoning (for a bit of a summary see my bio). Which of course makes it very hard for the people involved to understand why their systems fail at it; or, often, that they fail at it.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#309
post #230

Earlier quoted context omitted.

> In this particular case, is there any reason why we simply can't take their word for it? My concern is that if we take their word for it, we're actually buying into two assumptions which (AFAIK) are both unproven: 1. That "Internal Monologues" (not consciously forced by attention) exist in the first place, as opposed to being false-memories generated after-the-fact by our brain to explain/document a non-language pr…

Not only are they unproven, but are ultimately not provable at all. Some people will say yes, some people will say no. Probably we can take their word for it, but in the simplest case they could just lie (in either direction) and we would have no way to tell. In short, maybe these inner monologues exist and maybe they don't, but science can't comment on that. That said, it is clearly something we are interested in, b…

> Probably we can take their word for it, but in the simplest case they could just lie (in either direction) and we would have no way to tell.

Individually, no, but in general, for people to consistently lie about this particular thing at scale would be extremely unusual, given that people rarely lie if there's no reason for it. Going by this baseline, you could assume upward of 50% of replies are honest (even if mistaken), otherwise you'd have to explain why do you believe people would suddenly lie about that particular thing.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#310

Earlier quoted context omitted.

Why are so many people so insistent on saying this? I’m guessing you are in denial that we can make a simulated reasoning machine?

People keep saying it because that's literally how LLMs work. They run Montecarlo sampling over a very impressive latent linguistic space. These models are not fundamentally different than the Markov chains of yore except that these latent representations are incredibly powerful. We haven't even started to approach the largest problem which is moving beyond what is essentially a greedy token level search of this ling…

> LLMs are not reasoning machines. They are basically semantic compression machines with a build in search feature.

This is just a god of the gaps argument. Understanding is a form of semantic compression. So you're saying we have a system that can learn and construct a database of semantic information, then search it and compose novel, structured and coherent semantic content to respond to an a priori unknown prompt. Sounds like a form of reasoning to me. Maybe it's a limited deeply flawed type of reasoning, not that human reason is perfect, but that doesn't support your contention that it's not reasoning at all.

Post reply on HN