The popular retelling of how Einstein created Special Relativity to "Resolve the contradictions of Michelson-Morly experiments" is very reductive to the history of the question. The epitome is the quote from the paper: > From the two postulates, Einstein derived the Lorentz trans- formation ... If Einstein derived them, who is "Lorentz"? The groundwork for Special Relativity was the study of electrodynamics and symme…
Einstein seems to be so conveniently dismissive that he knew about the seminal Michelson-Morley experiments but he probably knew about it too well [1]. Einstein is not the first great scientist who are in denial of other important prior contributions, and he also not the last one. Newton also probably knew too well about Al-Haytham (Alhazen), arguably the father of modern science, and his breakthrough experiments but…
Position: LLMs Can't Jump
121–130 of 233 posts
Re: Position: LLMs Can't Jump
#122I had a related insight, but in the domain of humor [1]. LLMs are inherently probabilistic, and there's currently no mechanism for producing an orthogonal directional change in the path traced through a latent space which is also contextually relevant (landing on a punch line). In other words, LLMs are fundamentally incapable of making intuitive/orthogonal leaps in context. It might be possible to add this capability…
"Probabilistic next word prediction" and "humor" sit about as far apart as "modulating airflow with meat flaps" and "humor" do. One is an interface through which an action is performed and the other is a highly abstract capability.
Would you claim that a podcast comedian is fundamentally incapable of being funny because all he ever does is wiggle the air with his throat meat flaps? Probably not.
Absolutely nothing about "probabilistic next word prediction" forbids "making intuitive/orthogonal leaps in context". The interface is expressive enough.
And empirically? The "sense of humor" in LLMs is yet another "a function of model scale" capability. GPT-4.5 was reportedly funnier than both GPT-4o and o1. Fable 5 is reportedly funnier than Opus 4.x. It's one of those ever-elusive "big model smell" signs that are hard to measure with anything other than vibes.
Under the "humor as an opposed social intelligence test" family of hypothesis, what "being funny" reflects is the funny guy's ability to model and predict you and your reactions. For the comedian to be able to make the audience laugh, he must know his audience well, model it accurately enough to be able to spot the "breaking points" of humor, things they'd find unexpected and clever and thus "funny", and then weave those things into the jokes.
Then, a bigger LLM gets better at humor because it has a more accurate model of how humans think of things - including the "ha-ha" gaps. It's a "theory of mind" capability. It's not "special", it's just hard.
Re: Position: LLMs Can't Jump
#123Re: Position: LLMs Can't Jump
#124It's actually possible to answer this question rigorously:
1. Define a scientific result which qualifies as a "jump". They should be frequent enough that they happen every year - otherwise one might say humans can't jump either.
2. Identify all such "jumps" in articles published in 2026, and use LLM with 2025 knowledge cut-off to re-derive these results with minimal amount of information.
It really irks me that people boost these low-effort articles just because they confirm pre-conceived notion that LLMs are limited
Re: Position: LLMs Can't Jump
#125Maybe modulating temperature can help here: have the LLM come up with ideas at high temperature, and then critique them at low. This is also tied to halucinations: it is something that humans do (for writing fiction, and for "jumps") - but what LLMs currently lack is intellectual honesty. Coming up with bullshit is fine (and in this context valuable) - the important bit is putting those ideas through some form of rig…
I wonder if giving the models context of the temperature of its past generations would help here. Like a thinking mode that deliberately has a section that is high temperature, while the rest is lower.
If we could make progress in that area, maybe CoT could gradually decrease as it approaches its limit, or maybe the LLM could control the temperature of the next token itself (how this would be trained, I have no idea).
Re: Position: LLMs Can't Jump
#126Earlier quoted context omitted.
I think this can't work because an LLM needs too much data, and before the internet there probably just wasn't enough to get close to what we have now
Why couldn't an LLM, if it was smart enough, generate and consume its own data? I know the answer: because it leads to model collapse. But why is that? Wouldn't a smart model not collapse? It's seeming like they keep getting smarter because we keep pouring more of our own knowledge into them, not because they are actually getting smarter. And yes, sometimes a dumb but persistent bruteforcer can make new discoveries.
Re: Position: LLMs Can't Jump
#127This is literally an opinion of one dude which is not backed by any kind of quantitative evidence. It's actually possible to answer this question rigorously: 1. Define a scientific result which qualifies as a "jump". They should be frequent enough that they happen every year - otherwise one might say humans can't jump either. 2. Identify all such "jumps" in articles published in 2026, and use LLM with 2025 knowledge…
Re: Position: LLMs Can't Jump
#128I had a related insight, but in the domain of humor [1]. LLMs are inherently probabilistic, and there's currently no mechanism for producing an orthogonal directional change in the path traced through a latent space which is also contextually relevant (landing on a punch line). In other words, LLMs are fundamentally incapable of making intuitive/orthogonal leaps in context. It might be possible to add this capability…
You have committed the classic blunder of confusing your abstraction layers. "Probabilistic next word prediction" and "humor" sit about as far apart as "modulating airflow with meat flaps" and "humor" do. One is an interface through which an action is performed and the other is a highly abstract capability. Would you claim that a podcast comedian is fundamentally incapable of being funny because all he ever does is w…
I think you're saying that you can eventually train models to arrive at that destination by training on existing jokes, effectively encoding these leaps as probabilities.
In that case, the model isn't actually making an intuitive/comedic leap; they're just following new probability chains in attempting to approximate examples they've seen in training.
I'm suggesting that something architecturally different is necessary to create a model which can make intuitive/comedic leaps.
Try to get a frontier model to write a clever, funny joke which hasn't been seen before. Or, try to get it to make an intuitive leap that leads to a novel discovery.
You can use them to guide your own efforts along these lines, as a sounding board. But with current architecture I just don't think either is possible for an LLM to do on its own.
Re: Position: LLMs Can't Jump
#129Earlier quoted context omitted.
Was the ethical requirement to discuss and acknowledge predecessors work as developed in Newton's time as in our own? I would naturally expect not because it would seem to me to be the kind of thing that develops over time, but I could be wrong as I am not a historian of science.
John Maynard Keynes is famously quoted as stating: “[Newton] was not the first of the age of reason. He was the last of the magicians.”