Live data from Hacker News

Andrej Karpathy – It will take a decade to work through the issues with agents

dwarkesh.com

861–870 of 1001 posts

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#861
post #835

Earlier quoted context omitted.

This world model talk is interesting, and Yann Lecunn has broached on the same topic, but the fact is there are video diffusion models that are quite good at representing the "video world" and even counterfactually and temporally coherently generating a representation of that "world" under different perturbations. In fact you can go to a SOTA LLM today, and it will do quite well at predicting the outcomes of basic co…

> It's incredibly difficult to compress information without have at least some internal model of that information. Whether that model is a "world model" that fits the definition of folks like Sutton and LeCunn is semantic. Sutton's emphasizes his point by saying is that LLMs trying to reach AGI is futile because their world models are less capable that a squirrel's, in part because the squirrel has direct experiences…

Except Sutton has no idea or even a clue about the internal model of a squirrel. He just uses it as a symbol for utterly stupid but still smarter than an LLM. It’s semantic manipulation in attempt to prove his point but he proves nothing.

We have no idea how much of the world a squirrel understands. We understand LLMs more than squirrels. Arguably we don’t know if LLMs are more intelligent than squirrels.

> Finally he says if you could recreate the intelligence of a squirrel you'd be most of the way toward AGI, but you can't do that with an LLM.

Again he doesn’t even have a quantitative baseline for what intelligence means for a squirrel and how intelligent a squirrel is compared to an LLM. We literally have no idea if LLMs are more intelligent or less and no direct means of comparing what is more or less an apple and an orange.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#862
post #68

>What takes the long amount of time and the way to think about it is that it’s a march of nines. Every single nine is a constant amount of work. Every single nine is the same amount of work. When you get a demo and something works 90% of the time, that’s just the first nine. Then you need the second nine, a third nine, a fourth nine, a fifth nine. While I was at Tesla for five years or so, we went through maybe three…

In my experience with AI it's steeper than that: the jump from 90% to 99% is much harder than the jump from 0 to 90%

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#863

Earlier quoted context omitted.

And that last 5% is the toughest nut to crack. There is a reason waymo is way ahead even if they can not scale. Cameras are passive devices with relatively poor dynamic range and low light behavior. They are nowhere near a match/replacement for the human eye. Just try to picture a 5 year old at dusk or indoors and what you see will not be what you get.

Agree that the last fiew percentage points are exponentially more difficult each step of the way. What's your metric for saying Waymo is ahead, in terms of tech? They are strictly geo fenced, limited to specific road types, and often get stuck/confused. Also their system is very expensive, and not scalable to million of cars. Your point about cameras seems odd. Cameras have much better low light performance than huma…

waymo already has driverless taxi service in a major us city and is expanding. Tesla is in the process. again this is if they cover the last 5%. Scalability arguments wont matter when they can not launch such a service. And no, cmos cameras are close but are not better than the human eye in low light unless you have an ir camera and can flood everywhere with active ir lights. they are certainly inferior in dynamic range. I have been doing vision for more than two decades and I would not be comfortable in a camera only robotaxi at high speed. Certainly not at night or under adverse weather conditions. But this is all speculation of course. Considering fully autonomous driving at scale has been a major unrealised promise for the past 10 years, I stand by my assessment until I see a major advancement in camera technology or affordable active sensors.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#864
post #68

>What takes the long amount of time and the way to think about it is that it’s a march of nines. Every single nine is a constant amount of work. Every single nine is the same amount of work. When you get a demo and something works 90% of the time, that’s just the first nine. Then you need the second nine, a third nine, a fourth nine, a fifth nine. While I was at Tesla for five years or so, we went through maybe three…

[dead]

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#865

Earlier quoted context omitted.

This world model talk is interesting, and Yann Lecunn has broached on the same topic, but the fact is there are video diffusion models that are quite good at representing the "video world" and even counterfactually and temporally coherently generating a representation of that "world" under different perturbations. In fact you can go to a SOTA LLM today, and it will do quite well at predicting the outcomes of basic co…

Photons hit a human eye and then the human came up with language to describe that and then encoded the language into the LLM. The LLM can capture some of this relationship, but the LLM is not sensing actual photons, nor experiencing actual light cone stimulation, nor generating thoughts. Its "world model" is several degrees removed from the real world. So whatever fragment of a model it gains through learning to comp…

This comment is hallucinatory in nature as it is in direct conflict with the in the ground reality of LLMs.

The LLM has both light (aka photons) and language encoded into its very core. It is not just language. You seemed to have missed the boat with all the ai generated visuals and videos that are now inundating the internet.

Your flawed logic is essentially that LLMs are unable to model the real world because they don’t encode photonic data into the model. Instead you think they only encode language data which is an incredibly lossy description of reality. And this line of logic flies against the ground truth reality of the fact that LLMs ARE trained with video and pictures which are essentially photons encoded into data.

So what should be the proper conclusion? Well look at the generated visual output of LLMs. These models can generate video that is highly convincing and often with flaws as well but often these videos are indistinguishable from reality. That means the models have very well done but flawed simulations of reality.

In fact those videos demonstrate that LLMs have extremely high causal understanding of reality. They know cause and effect it’s just the understanding is imperfect. They understand like 85 percent of it. Just look at those videos of penguins on trampolines. The LLM understands what happens as an effect after a penguin jumps on a trampoline but sometimes an extra penguin teleports in which shows that the understanding is high but not fully accurate or complete.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#866

Agency. If one studied the humanities they’d know how incredible a proposal “agentic” AI is. In the natural world, agency is a consequence of death: by dying, the feedback loop closes in a powerful way. The notion of casual agency (I’m thinking of Jensen Huang’s generative > agentic > robotic insistence) is bonkers. Some things are not easily speedrunned. (I did listen to a sizable portion of this podcast while makin…

Models die too - the less agentic ones are out-competed by the more agentic ones. Every AI lab brags how "more agentic" their latest model is compared to the previous one and the competition, and everybody switches to the new model.

Yes but the point is that models must be imminently aware of their impending death to force the calculation of tradeoffs.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#867

Earlier quoted context omitted.

Photons hit a human eye and then the human came up with language to describe that and then encoded the language into the LLM. The LLM can capture some of this relationship, but the LLM is not sensing actual photons, nor experiencing actual light cone stimulation, nor generating thoughts. Its "world model" is several degrees removed from the real world. So whatever fragment of a model it gains through learning to comp…

I agree with this. A metaphor I like is that the reason why humans say the night sky is beautiful is because they see that it is, whereas an LLM says it because it’s been said enough times in its training data.

Guys you realize that you can go to ChatGPT right now and it can generate an actual picture of the night sky because it has seen thousands of pictures and drawings of the actual night sky right?

Your logic is flawed because your knowledge is outdated. LLMs are encoding visual data, not just “language” data.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#868
post #786

Earlier quoted context omitted.

Agreed. I'd also add he's intellectually honest enough to not overhype what's happening just to hype whatever he's working on or appear to be a thought leader. Just very clear, pragmatic, and intellectually honest thought about the reality of things.

It's almost like having more money than you'll ever know what to do with lets you say and do what you _actually_ want to do.

Most people don’t take this opportunity, though.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#869

Maybe I'm being too simplistic, but I think we're mixing two distinct debates. Today we have an extraordinary invention—comparable to the wheel in its time. That invention is: predictive inference over all human knowledge. Period. I don't like calling it "Artificial Intelligence" because it's not intelligence; it's a prediction system that can project responses by illuminating patterns across all human knowledge enca…

I think this comparison is all wrong. The internet is more closer to the notion of a wheel - the internet has done amazing stuff just as the wheel has and nobody foresaw the impact the internet would have and how the underlying technologies that power it have evolved. Just like how a wheel moves stuff, the internet is the medium through which bits are transmitted and received.

Thanks for your view! My analogy was intentional—I wanted to talk about revolutionary tools that extend human capabilities, not about the foundational infrastructure itself. Of course, the internet is a fundamental platform like the wheel, but I’m focusing on what’s built on top of that—how new tools like predictive inference change the landscape again. Analogies can work at different layers. I just chose the tool, not the medium.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#870
post #835

Earlier quoted context omitted.

> It's incredibly difficult to compress information without have at least some internal model of that information. Whether that model is a "world model" that fits the definition of folks like Sutton and LeCunn is semantic. Sutton's emphasizes his point by saying is that LLMs trying to reach AGI is futile because their world models are less capable that a squirrel's, in part because the squirrel has direct experiences…

Except Sutton has no idea or even a clue about the internal model of a squirrel. He just uses it as a symbol for utterly stupid but still smarter than an LLM. It’s semantic manipulation in attempt to prove his point but he proves nothing. We have no idea how much of the world a squirrel understands. We understand LLMs more than squirrels. Arguably we don’t know if LLMs are more intelligent than squirrels. > Finally h…

> We have no idea how much of the world I squirrel understands. We understand LLMs more than squirrels

Based on our understanding of biology and evolution we know that a squirrel brain works more similarly to the way we humans do vs an LLM.

To the extent we understand LLMs, it's because they are strictly less complex than both ours and squirrels' brains, not because they are better model for our intelligence. They are a thin simulation of human language generation capability mediated via text.

We also see that a squirrel, like us, is capable of continuous learning driven by its own goals, all on an energy budget many orders of magnitude lower than LLMs. That last part is a strong empirical indication that suggests that LLMs are a dead end for AGI, given that the real world employs harsh energy constraints on biological intelligences.

Also remember that Sutton is still of an AI maximalist. He isn't saying that AGI isn't possible, just that LLMs can't get us there.

Post reply on HN