Live data from Hacker News

What can LLMs never do?

strangeloopcanon.com

81–90 of 385 posts

Re: What can LLMs never do?

#81
post #67

Earlier quoted context omitted.

If you're contending that LLMs are incapable of reasoning, you're saying that there's no reasoning task that an LLM can do. Is that what you're saying? Because I can easily find an example to prove you wrong.

It could be that all reasoning displayed is showing existing information - so there would be no reasoning, but that aside, what I meant is being able to reason in any consistent way. Like a machine that only sometimes gets an addition right isn't really capable of addition.

The former is easy to test, just make up your own puzzles and see if it can solve them.

"Incapable of reasoning" doesn't mean "only solves some logic puzzles". Hell, GPT-4 is better at reasoning than a large number of people. Would you say that a good percentage of humans are poor at reasoning too?

Re: What can LLMs never do?

#82
post #78

If we're trying to quantify what they can NEVER do, I think we'd have to resort to some theoretical results rather than a list empirical evidence of what they can't do now. The terminology you'd look for in the literature would be "expressibility". For a review of this topic, I'd suggest: https://nessie.ilab.sztaki.hu/~kornai/2023/Hopf/Resources/st... The authors of this review have themselves written several article…

Thank you for sharing this here. Rigorous work on the "expressibility" of current LLMs (i.e., which classes of problems can they tackle?) is surely more important , but I suspect it will go over head of most HN readers, many of whom have minimal to zero formal training on topics relating to computational complexity.

Yes, but unfortunately that doesn't answer the question the title poses.

Re: What can LLMs never do?

#83
post #81

Earlier quoted context omitted.

It could be that all reasoning displayed is showing existing information - so there would be no reasoning, but that aside, what I meant is being able to reason in any consistent way. Like a machine that only sometimes gets an addition right isn't really capable of addition.

The former is easy to test, just make up your own puzzles and see if it can solve them. "Incapable of reasoning" doesn't mean "only solves some logic puzzles". Hell, GPT-4 is better at reasoning than a large number of people. Would you say that a good percentage of humans are poor at reasoning too?

Not just logic puzzles but also applying information, and, yes, I tried a few things.

People/humans tend to be pretty poor, too (training can help, though), as it isn't easy to really think through and solve things - we don't have a general recipe to follow there and neither do LLMs it seems (otherwise it shouldn't fail).

What I am getting at is that as far as a reasoning machine is concerned, I'd want it to be like a pocket calculator is for arithmetic, i.e., it doesn't fail other than in some rare exceptions - and not inheriting human weaknesses there.

Re: What can LLMs never do?

#84
post #78

Earlier quoted context omitted.

Thank you for sharing this here. Rigorous work on the "expressibility" of current LLMs (i.e., which classes of problems can they tackle?) is surely more important , but I suspect it will go over head of most HN readers, many of whom have minimal to zero formal training on topics relating to computational complexity.

Yes, but unfortunately that doesn't answer the question the title poses.

The OP is not trying to answer the question. Rather, the OP is asking the question and sharing some thoughts on the motivations for asking it.

Re: What can LLMs never do?

#85
Some of these "never do" things are just artifacts of textual representation, and if you transformed wordl/sudoku into a different domain it would have a much higher success rate using the exact same transformer architecture.

We don't need to create custom AGI for every domain, we just need a model/tool catalog and an agent that is able to reason well enough to decompose problems into parts that can be farmed out to specialized tools then reassembled to form an answer.

Re: What can LLMs never do?

#86
post #24
post #19

> But then I started asking myself how can we figure out the limits of its ability to reason Third paragraph. The entire article is based on the premise LLMs are supposed to reason, which is wrong. They don't, they're tools to generate text.

I really hate this reductive, facile, "um akshually" take. If the text that the text-generating tool generates contains reasoning, then the text generation tool can be said to be reasoning, can't it. That's like saying "humans aren't supposed to reason, they're supposed to make sounds with their mouths".

At some point if you need to generate better text you need to start creating a model of how the world works along with some amount of reasoning. The "it's just a token generator" argument fails to get this part. That being said I don't think just scaling LLMs are going to get us AGI but I don't have any real arguments to support that

Re: What can LLMs never do?

#87

If we're trying to quantify what they can NEVER do, I think we'd have to resort to some theoretical results rather than a list empirical evidence of what they can't do now. The terminology you'd look for in the literature would be "expressibility". For a review of this topic, I'd suggest: https://nessie.ilab.sztaki.hu/~kornai/2023/Hopf/Resources/st... The authors of this review have themselves written several article…

We have to be a bit more honest about the things we can actually do ourselves. Most people I know would flunk most of the benchmarks we use to evaluate LLMs. Not just a little bit but more like completely and utterly and embarrassingly so. It's not even close; or fair. People are surprisingly alright at a narrow set of problems. Particularly when it doesn't involve knowledge. Most people also suck at reasoning (unless they had years of training), they suck at factual knowledge, they aren't half bad at visual and spatial reasoning, and fairly gullible otherwise.

Anyway, this list looks more like a "hold my beer" moment for AI researchers than any fundamental objections for AIs to stop evolving any further. Sure there are weaknesses, and paths to address those. Anyone claiming that this is the end of the road in terms of progress is going to be in for some disappointing reality check probably a lot sooner than is comfortable.

And of course by narrowing it to just LLMs, the authors have a bit of an escape hatch because they conveniently exclude any further architectures, alternate strategies, improvements, that might otherwise overcome the identified current weaknesses. But that's an artificial constraint that has no real world value; because of course AI researchers are already looking beyond the current state of the art. Why wouldn't they.

Re: What can LLMs never do?

#88
post #40

Earlier quoted context omitted.

> If the text that the text-generating tool generates contains reasoning, then the text generation tool can be said to be reasoning, can't it. I don't know... you're still describing a talking parrot here, if you'd ask me.

What's the difference between a human and a talking parrot that can answer any question you ask it?

Can any question be answered? As long as any reaction on a question is considered an answer, then I see no difference between a human and a parrot.

Re: What can LLMs never do?

#89

If we're trying to quantify what they can NEVER do, I think we'd have to resort to some theoretical results rather than a list empirical evidence of what they can't do now. The terminology you'd look for in the literature would be "expressibility". For a review of this topic, I'd suggest: https://nessie.ilab.sztaki.hu/~kornai/2023/Hopf/Resources/st... The authors of this review have themselves written several article…

This is also a good paper on the subject:

What Algorithms can Transformers Learn? A Study in Length Generalization https://arxiv.org/abs/2310.16028

Re: What can LLMs never do?

#90
post #62

I have been trying to generate some text recently using the ChatGPT API. No matter how I word “Include any interesting facts or anecdotes without commenting on the fact being interesting” it ALWAYS starts out “One interesting fact about” or similar phrasing. I have honestly spent multiple hours trying to word the prompt so it will stop including introductory phrases and just include the fact straight. I have gone so…

Have you tried a simple "No pretext or posttext, return the result in a code block"?
Post reply on HN