Live data from Hacker News

Position: LLMs Can't Jump

openreview.net

41–50 of 233 posts

Re: Position: LLMs Can't Jump

#41

It's a very shaky position, and the empirical track record of "LLMs can't..." is in itself a reason to call it into doubt. Every "can't" of this nature was followed by a discovery of "they can, just poorly", and then by that "poorly" improving steadily generation to generation. The paper doesn't provide a way to measure or quantify this elusive "jumping" capability, not even as an approximation. It just throws "can't…

> the empirical track record of "LLMs can't..." is in itself a reason to call it into doubt ... Every "can't" of this nature was followed by a discovery of "they can, just poorly", and then etc

What LLMs are "fundamentally incapable" of doing has striking parallels to https://en.wikipedia.org/wiki/God_of_the_gaps

Re: Position: LLMs Can't Jump

#42
post #21

I think this is more of a function of the harness and the environment than the LLM. I've seen some LLM interactions over complex environments like Godot and Unity that challenges the notion that there is no "jumping" going on at all. An LLM in isolation from its environment might as well be a brain in a vat in some dark cave. You need an external environment to sample from and act upon to make forward progress.

Why does the cave need to be dark if its just a brain in a vat?

'Cause the brain is in a glass vat. and brains have eyes too. https://www.youtube.com/watch?v=pc3uWxs9adw

Re: Position: LLMs Can't Jump

#43
post #22

The theory is that creative leaps in theoretical physics require a grounding in sensory experience, but the obvious counter-argument is that humans can make creative leaps in abstract fields without such sensory grounding. They do address this at the end, saying "In abstract domains such as Mathematics or Computer Science, the Sense Experience (E) may be grounded in high-dimensional topology or have other goals such…

> ...but the obvious counter-argument is that humans can make creative leaps in abstract fields without such sensory grounding... But we have no idea at how good humans are at that. Given the appalling failures of humans to handle even basic statistical situations like identifying that the same thing happens over and over, it might be that they are hilariously bad at creative leaps in abstract fields, it is just we h…

Most humans are probably bad at it. Some humans are very good at it. I'd even argue most science does not demand these kinds of leaps and is mostly concerned with incremental improvements, or proving or disproving other people's abductions

The case study they chose is literally Albert Einstein coming up with General Relativity, something most scientists of his time were not able to do

Re: Position: LLMs Can't Jump

#44
We just put different things together and then we evaluate it.

In math its simple: does the verification say its okay.

If its mechanical: is any property better than what we have already.

etc.

Re: Position: LLMs Can't Jump

#45

Came for: "A computer once beat me at chess, but it was no match for me at kick boxing." TFA was actually about leaps of intuition, sadly. One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself.

> One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself.

Could be, but preventing leakage from more modern stuff can be challenging.

This was attempted with Victorian public domain content: https://www.estragon.news/mr-chatterbox-or-the-modern-promet...

I can't find the citation right now, but I think people found it was leaking anachronisms? So this probably wasn't as well filtered as the creator had hoped?

Re: Position: LLMs Can't Jump

#46

Worth reposting a follow-up tweet from the author Tom Zahavy [1] after this made the rounds on X/Twitter recently: > A few reflections on my "LLMs Can’t Jump" paper: > My position paper recently got some traction here, so I wanted to share a few thoughts and clarify a few things. > First things first: some people are framing this as "DeepMind is throwing cold water on AI for science" or claiming the paper argues LLMs…

Moreover anyone glomming onto this paper for goal-post-shifting “AI can never” should:

1. Read the last sentence of the abstract, and

2. Reflect that frontier reasoning agents already increasingly integrate multimodal models.

Re: Position: LLMs Can't Jump

#47

Worth reposting a follow-up tweet from the author Tom Zahavy [1] after this made the rounds on X/Twitter recently: > A few reflections on my "LLMs Can’t Jump" paper: > My position paper recently got some traction here, so I wanted to share a few thoughts and clarify a few things. > First things first: some people are framing this as "DeepMind is throwing cold water on AI for science" or claiming the paper argues LLMs…

I mean Einstein had help, he was networked with the best scientific minds of the planet and his discoveries were grounded in experimental results that contradicted existing theories (at least for specialized relativity), and without Riemann’s work he wouldn’t have been able to formulate his theory either. So not sure if AI couldn’t do that if you kept feeding it with new research results and let it correspond with top human scientists. Einstein was a genius but I don’t think his thought process is beyond what an LLM could simulate. And again this is probably the most impressive scientific achievement in theoretical physics in the 20. century so maybe it’s hanging the bar a bit high for LLMs.

Re: Position: LLMs Can't Jump

#50
post #36

Came for: "A computer once beat me at chess, but it was no match for me at kick boxing." TFA was actually about leaps of intuition, sadly. One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself.

Is that how chessboxing was invented? Genuinely asking.

Nope.

Chessboxing was the invention of comics book artist Enki Bilal (and he's credited with this in Wikipedia). I first saw it in his Nikopol trilogy. Because life is weird, it then became a real thing.

It's unrelated to computers playing chess. It predates Kasparov's first defeat by Deep Blue. I don't remember any mention of computers being good at chess in the trilogy, either. Or any computers, for that matter.

Post reply on HN