Live data from Hacker News

Andrej Karpathy – It will take a decade to work through the issues with agents

dwarkesh.com

971–980 of 1001 posts

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#971

Earlier quoted context omitted.

I don't think he is saying agents are not useful at all, just that they are not anywhere near the capability of human software developers. Karpathy later says he used agents to write the Rust translation of algorithms he wrote in Python. He also explicitly says that agents can be useful for writing boilerplate or for code that can be very commonly found online. So I don't think he is saying they are not useful at all…

I’m not saying he’s saying agents aren’t useful at all. It’s literally in the quotes I provided that he says they are useful for some subset of tasks. I’m saying that he is answering the question “are agents useful at all”. not “can agents replace humans”. His answer is mostly not. He generally prefers autocomplete. But they are useful for some limited tasks.

Saying "He’s talking about whether agents are currently useful at all" is negatively loaded. It is very easy to take that and assume the answer is "no" based on the "at all".

If you wanted to be more neutral, you could have said something like "He's also questioning how useful agents really are today". That wouldn't have implied that they're not useful at all, but instead that they're less useful than people are claiming.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#972

Earlier quoted context omitted.

> We always had the math to show that scale wasn't enough Math, to show that scale (presumably of LLMs) wasn't enough for AGI? This sounds like it would be quite a big deal, what math is that?

As someone who is invested in researching said math, I can say with some confidence that it does not exist, or at least not in the form claimed here. That's the whole problem. I would be ecstatic if it did though, so if anyone has any examples or rebuttal, I would very much appreciate it.

Let me clarify. I was too vague and definitely did not express things accurately. That is on me.

We have the math to show that it can be impossible to distinguish two explanations through data processing alone. We have examples of this in science, a long history of it in fact. Fundamentally there is so much that we cannot conclude from processing data alone. Science (the search of knowledge) is active. It doesn't require just processing existing data, it requires the search for new data. We propose competing hypotheses that are indistinguishable from the current data and seek out the data which distinguishes them (a pain point for many of the TOEs like String Theory). We know that data processing alone is insufficient for explanation. We know it cannot distinguish confounders. We know it cannot distinguish causal graphs (e.g. distinguish triangular maps. We are able to create them, but not distinguish them through data processing alone). The problem with scaling alone is that it makes the assertions that data processing is enough. Yet we have so much work (and history) telling us that data processing is insufficient.

The scaling math itself also shows a drastic decline in performance with scale and often do not suggest convergence even with infinite data. They are power laws with positive concavity, requiring exponential increase in data and parameters for marginal improvements on test loss. I'm not claiming that we need zero test loss to reach AGI, but the results do tell us that if this is strongly correlated then we'll need to spend an exponential amount more to achieve AGI even if we are close. By our measures, scaling is not enough unless we are sufficiently close. Even our empirical results align with this as despite many claiming that scale is all we need, we are making significant changes to the model architectures and training procedures (including optimizers). We are making these large changes because throwing the new data at the old models (even when simply increasing the number of parameters) does not work out. It is not just the practicality, it is the results. The scaling claim has always been a myth used to drive investments since it is a nice simple story that says that we can get there by doing what we've already been doing, just more. We all know that these new LLMs aren't dramatic improvements off their previous versions, despite being much larger, more efficient, and having processed far more data.

[side note]: We even have my namesake who would argue that there are truths which are not provably true with a system that is both consistent and efficient (effectively calculable). But we need not go that far, as omniscience is not a requirement for AGI. Though it is worth noting for the limits of our models, since at the core this matters. Changing our axioms changes the results, even with the same data. But science doesn't exclusively use a formal system, nor does it use a single one.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#973

Earlier quoted context omitted.

As someone who is invested in researching said math, I can say with some confidence that it does not exist, or at least not in the form claimed here. That's the whole problem. I would be ecstatic if it did though, so if anyone has any examples or rebuttal, I would very much appreciate it.

You're right that there is no purely mathematical argument; it's almost non-sensical to claim such. Instead you can simply make the rather reasonable observation that LLMs are a product of their training distribution, which only contains partial coverage of all possible observable states of the world. Some highly regular observable states are thus likely missing, but an embodied agent (like a human) would be able to…

My claim is more about that data processing is not enough. I was too vague and I definitely did not convey myself accurately. I tried to clarify a bit in a sibling comment to yours but I'm still unsure if it is sufficient tbh.

For embodiment, I think this is sufficient but not necessary. A key part to the limitation is that the agent cannot interact with its environment. This is a necessary feature for distinguishing competing explanations. I believe we are actually in agreement here, but I do think we need to be careful how we define embodiment. Because even a toaster can be considered a robot. It seems hard to determine what does not qualify as a body when we get to the itty gritty. But I think in general when people are talking about embodiment they are discussing the capability of being interventional.

By your elaboration I believe we agree since part of what I believe to be necessary is the ability to self-analyze (meta-cognition) to determine low density regions of its model and then to be able to seek out and rectify this (intervention). Data processing is not sufficient for either of those conditions.

Your prompt is, imo, more about world modeling, though I do think this is related. I asked Claude Sonnet 4.5 with extended thinking enabled and it also placed itself outside the room. Opus 4.1 (again with extended thinking), got the answer right. (I don't use a standard system prompt, though that is mostly to make it not syncopathic and to try to get it to ask questions when uncertain and enforce step by step thinking)

  From the perspective of the person in the room, your right arm would be on their right side as you walk out.
  
  Here's why: When you initially walk into the room facing the person, your right arm appears on their left side (since you're facing each other). But when you turn around 180 degrees to walk back out, your back is now toward them. Your right arm stays on your right side, but from their perspective it has shifted to their right side.

  Think of it this way - when two people face each other, their right sides are on opposite sides. But when one person turns their back, both people's right sides are now on the same side.
The CoT output is a bit more interesting[0]. Disabling my system prompt gives an almost identical answer fwiw. But Sonnet got it right. I repeated the test in incognito after deleting the previous prompts and it continued to get it right, independent of my system prompt or extended thinking.

I don't think this proves a world model though. Misses are more important than hits, just as counter examples are more important than examples in any evidence or proof setting. But fwiw I also frequently ask these models variations on river crossing problems and the results are very shabby. A few appear spoiled now but they are not very robust to variation and that I think is critical.

I think an interesting variation of your puzzle is as follows

  Imagine you walked into a room through a doorway. Then you immediately turn around and walk back out of the room. 

  From the perspective of a person in the room, facing the door, which side would your right arm be? Please explain.
I think Claude (Sonnet) shows some subtle but important results in how it answers

  Your right arm would be on their right side.
  When you turn around to walk back out, you're facing the same direction as the person in the room (both facing the door). Since you're both oriented the same way, your right side and their right side are on the same side.
This makes me suspect there's some overfitting. CoT correctly uses "I"[1].

It definitely isn't robust to red herrings[2], and I think that's a kicker here. It is similar to failure results I see in any of these puzzles. They are quite easy to break with small variations. And we do need to remember that these are models trained on the entire internet (including HN comments), so we can't presume this is a unique puzzle.

[0] http://0x0.st/K158.txt

[1] http://0x0.st/K15T.txt

[2] http://0x0.st/K15m.txt

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#974

Earlier quoted context omitted.

They = Waymo

Karpathy talked about Waymo, and he said they aren't there yet. They still have humans in the loop via telemetry and there are parts of cities they won't go to.

Just go ride in one and make the decision for yourself. By any reasonable person's definition they are self driving.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#975

Earlier quoted context omitted.

Humans adapt and become more nines the more they learn about something. Humans also are liable in a lawful sense. This is a huge factor in any AI use case.

So, it's not really nines, but the lack of continuous learning and legal issues.

Continuous learning seems like one of real criteria of AGI

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#976

Earlier quoted context omitted.

This world model talk is interesting, and Yann Lecunn has broached on the same topic, but the fact is there are video diffusion models that are quite good at representing the "video world" and even counterfactually and temporally coherently generating a representation of that "world" under different perturbations. In fact you can go to a SOTA LLM today, and it will do quite well at predicting the outcomes of basic co…

Photons hit a human eye and then the human came up with language to describe that and then encoded the language into the LLM. The LLM can capture some of this relationship, but the LLM is not sensing actual photons, nor experiencing actual light cone stimulation, nor generating thoughts. Its "world model" is several degrees removed from the real world. So whatever fragment of a model it gains through learning to comp…

Entities equiped with two limited light sensitive captors encode through a network of carbon based chemical emitters a representation of what its flawed vision system manages to grasp biased towards self preservation.

What's the real world? I'm still puzzled by this reaction I see to LLM, not because I think LLM are undervalued, because most people seem to significantly overestimate what is human intelligence.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#977

Earlier quoted context omitted.

stop repeating that. first, it isn't true that intelligence is barely defined. https://arxiv.org/abs/0706.3639 second a definition is obviously not a prerequisite as evidenced by natural selection

> stop repeating that. first, it isn't true that intelligence is barely defined. https://arxiv.org/abs/0706.3639 I don't think he should stop, because I think he's right. We lack a definition of intelligence that doesn't do a lot of hand waving. You linked to a paper with 18 collective definitions, 35 psychologist definitions, and 18 ai researcher definitions of intelligence. And the conclusion of the paper was that…

cool, you can move goalposts and claim no true scotsman would define intelligence that way; in addition you are confusing sufficient and necessary.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#978

Earlier quoted context omitted.

stop repeating that. first, it isn't true that intelligence is barely defined. https://arxiv.org/abs/0706.3639 second a definition is obviously not a prerequisite as evidenced by natural selection

An Arxiv paper listing 70 different definitions of intelligence is not the evidence that you seem to think it is.

yes it is

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#979
>I feel like when I'm awake, I'm building up a context window of stuff that's happening during the day. But when I go to sleep, something magical happens where I don't think that context window stays around.

>There's some process of distillation into the weights of my brain. This happens during sleep and all this stuff. We don't have an equivalent [in LLMs] (23:09)

Seems to me that's one of the big things lacking in LLMs vs human thinking. People say LLMs can't lead on to AGI but that kind of thing is an avenue they could explore.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#980

Earlier quoted context omitted.

This world model talk is interesting, and Yann Lecunn has broached on the same topic, but the fact is there are video diffusion models that are quite good at representing the "video world" and even counterfactually and temporally coherently generating a representation of that "world" under different perturbations. In fact you can go to a SOTA LLM today, and it will do quite well at predicting the outcomes of basic co…

Photons hit a human eye and then the human came up with language to describe that and then encoded the language into the LLM. The LLM can capture some of this relationship, but the LLM is not sensing actual photons, nor experiencing actual light cone stimulation, nor generating thoughts. Its "world model" is several degrees removed from the real world. So whatever fragment of a model it gains through learning to comp…

Photons reflected off of objects are not the actual objects. I wouldn't go so far as to say that sensing these is a particularly special way to know about things compared to hearing or reading about them. Further, many humans do not sense photons yet seem to manage to have perfectly fine working world models.
Post reply on HN