Position: LLMs Can't Jump
31–40 of 233 posts
Re: Position: LLMs Can't Jump
#32You may think this is not a good test because an older (or say a smaller) LLM can study from the knowledge on the Internet and build. But we are like that - we can access the Universe through our senses.
Can we ever produce anything that is beyond this Universe? I think an LLM that is lacking in knowledge can build more complex systems as long as it can access more data.
Re: Position: LLMs Can't Jump
#33Every "can't" of this nature was followed by a discovery of "they can, just poorly", and then by that "poorly" improving steadily generation to generation.
The paper doesn't provide a way to measure or quantify this elusive "jumping" capability, not even as an approximation. It just throws "can't jump" out there, as if "abduction" is an established class of problem with known computational properties and requirements that the LLM architecture fails to satisfy. It's none of those things - and the paper makes the claim without backing it by anything but rhetoric attempts at persuasion.
The proposed solution is also dubious. The empirical track record of dedicated "world models" for reasoning and problem-solving is, frankly, downright abysmal. Even integrating multimodal data into LLMs has failed to yield general reasoning capability gains.
LeCun's misadventures in the field aside, the main frontier lab that pushes in favor of "improving reasoning via multimodal fusion" is GDM - and Gemini isn't exactly a paragon of frontier reasoning capabilities. It has strong multimodal capabilities, but lags behind both OpenAI and Anthropic in performance outside that - while Anthropic is the lab that always treated multimodal grounding as an afterthought, and still trades blows with OpenAI at the very edge of the performance frontier. Multimodal grounding seems to work great as a way to improve an AI's ability to deal with those specific modalities, but it falters outside that.
Now, it's not impossible that everyone who tried multimodal world models for reasoning is just doing it wrong, and there is an undiscovered recipe for multimodal grounding that results in a step change in AI capabilities. But the results we have so far suggest it to be unlikely.
Re: Position: LLMs Can't Jump
#34Re: Position: LLMs Can't Jump
#35I feel like you could just add some noise or randomness to the LLM and start approximating the leaps that the human mind uses to solve and understand unrelated things. Maybe that’s naive, it’s just coming from my organic computer in my skull.
Re: Position: LLMs Can't Jump
#36Came for: "A computer once beat me at chess, but it was no match for me at kick boxing." TFA was actually about leaps of intuition, sadly. One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself.
Re: Position: LLMs Can't Jump
#37Came for: "A computer once beat me at chess, but it was no match for me at kick boxing." TFA was actually about leaps of intuition, sadly. One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself.
I think this can't work because an LLM needs too much data, and before the internet there probably just wasn't enough to get close to what we have now
Re: Position: LLMs Can't Jump
#38Earlier quoted context omitted.
I think this can't work because an LLM needs too much data, and before the internet there probably just wasn't enough to get close to what we have now
Even simpler: Can GPT-2 anticipate and build Gwen/Deepseek? I think the answer is almost trivially "no", so I wonder what changed?
Re: Position: LLMs Can't Jump
#39And then such jump can be verified by a machine so human kind of plugs the intelligence gap.
That’s pretty exciting.
I always liked to provocatively call LLM „the new calculator”. Calculator for language.
We are so focused on creating a standalone intelligence that we didn’t notice how we massively augmented our own. That could be considered transhumanism holy grail if only interface brain-LLM was faster.
People need to understand that these things are tools. And every tool needs an operator to function. Tool doesn’t have its own goals, needs, wants or motives. It won’t do anything out of its own, it always exists in context of someone telling it what to do.
In light of that most of the panic and fear mongering is rather ridiculous. Calculator won’t replace you. It wont take over the world. It is just a tool.
You write a book with book generator? Cool, it can be used for this. We will judge output, not the methods. Sometimes we will judge people who have no taste in literature.