Live data from Hacker News

Andrej Karpathy – It will take a decade to work through the issues with agents

dwarkesh.com

431–440 of 1001 posts

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#431
post #226

Earlier quoted context omitted.

The interview which I've watched recently with Rich Sutton left me with the impression that AGI is not just a matter of adding more 9s. The interviewer had an idea that he took for granted: that to understand language you have to have a model of the world. LLMs seem to udnerstand language therefore they've trained a model of the world. Sutton rejected the premise immediately. He might be right in being skeptical here…

There is some evidence from Anthropic that LLMs do model the world. This paper[0] tracing their "thought" is fascinating. Basically an LLM translating across languages will "light up" (to use a rough fMRI equivalent) for the same concepts (e.g. bigness) across languages. It does have clusters of parameters that correlate with concepts, not just randomly "after X word tends to have Y word." Otherwise you would expect…

> Basically an LLM translating across languages will "light up" for the same concepts across languages

Which is exactly what they are trained to do. Translation models wouldn't be functional if they are unable to correlate an input to specific outputs. That some hiddel-layer neurons fire for the same concept shouldn't come as a surprise, and is a basic feature required for the core functionality.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#432

Earlier quoted context omitted.

In my view 'understand' is a folk psychology term that does not have a technical meaning. Like 'intelligent', 'beautiful', and 'interesting'. It usefully labels a basket of behaviors we see in others, and that is all it does. In this view, if a machine performs a task as well as a human, it understands it exactly as much as a human. There's no problem of how to do understanding, only how to do tasks. The 'problem' me…

Nonsense. A QC operator may be able to carry out a test with as much accuracy (or perhaps better accuracy, with enough practice) than the PhD quality chemist who developed it. They could plausibly do so with a high school education and not be able to explain the test in any detail. They do not understand the test in the same way as the chemist. If 'understand' is a meaningless term to someone who's spent 30 years in…

so your definition of "understand" is "able to develop the QC test (or explain tests already developed)"

I hate to break it to you, but the LLMs can already do all 3 tasks you outlined

It can be argued for all 3 actors in this example (the QC operator, the PhD chemist and the LLM) that they don't really "understand" anything and are iterating on pre-learned patterns in order to complete the tasks.

Even the ground-breaking chemist researcher developing a new test can be reduced to iterating on the memorized fundamentals of chemistry using a lot of compute (of the meat kind).

The mythical Understanding is just a form of "no true Scotsman"

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#433

5 decades. You have one decade to clean up your power use problem. If you don't you will find yourself in the next AI winter.

Power use is less important than model capability

AGI is either more scale or differing systems, or both

They can always optimize for power consumption after AGI has been reached

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#434
post #289

Earlier quoted context omitted.

>Why is there a presumption that we (as people who have only studied CS) know enough about biology/neuroscience/evolution to make these comparisons? Hubris.

The hubris here isn't CS people making comparisons, it's assuming biological substrate matters. Your brain is doing computation with neurotransmitters instead of transistors. So what? The "chemicals not electricity" distinction is pure carbon chauvinism, like insisting hydraulic computers can't be compared to electronic ones because water isn't electricity. Evolution didn't discover some mystical process that imbues…

> Your brain is doing computation with neurotransmitters instead of transistors.

This is an incredible simplification of the process and also just a small part of it. There is increasing evidence that quantum effects might play a part in the inner workings of the brain.

> Brains work despite being kludges of evolutionary baggage, not because biology unlocked some deeper truth about intelligence.

Now that is hubris.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#435

Earlier quoted context omitted.

Most likely because you'll be filthy reach from selling AGI and won't need to go after secondary revenue sources.

>Most likely because you'll be filthy reach from selling AGI Why? If AGI costs more than a human or operates slower than one, it may not be economical for people to buy it. By the time it becomes economical, competitors may have also cracked it reducing your ability to charge high margins on it.

Cost decreases with time

Humans can work on a problem 8 hours a day? You can run inference 24/7

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#436

Earlier quoted context omitted.

Nonsense. A QC operator may be able to carry out a test with as much accuracy (or perhaps better accuracy, with enough practice) than the PhD quality chemist who developed it. They could plausibly do so with a high school education and not be able to explain the test in any detail. They do not understand the test in the same way as the chemist. If 'understand' is a meaningless term to someone who's spent 30 years in…

> If 'understand' is a meaningless term to someone who's spent 30 years in AI research, I understand why LLMs are being sold and hyped in the way they are. I don't have quite as much time as robotresearcher, but I've heard their sentiment frequently. I've been to conferences, talked with people at the top of the field (I'm "junior", but published and have a PhD) where when asking deeper questions I'll get a frequent…

As a mathematician who also regularly publishes in these conferences, I am a little surprised to hear your take; your experience might be slightly different to mine.

Identifying limitations of LLMs in the context of "it's not AGI yet because X" is huge right now; it gets massive funding, taking away from other things like SciML and uncertainty analyses. I will agree that deep learning theory in the sense of foundational mathematical theory to develop internal understanding (with limited appeal to numerics) is in the roughest state it has even been in. My first impression there is that the toolbox has essentially run dry and we need something more to advance the field. My second impression is that empirical researchers in LLMs are mostly junior and significantly less critical of their own work and the work of others, but I digress.

I also disagree that we are disincentivised to find meaning behind the word "understanding" in the context of neural networks: if understanding is to build an internal world model, then quite a bit of work is going into that. Empirically, it would appear that they do, almost by necessity.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#437

I have massive respect for Andrej, my first encounter with "him" was following his tutorials/notes when he was a grad student/tutor for AI/ML. I was a lot disappointed when he went to work for Tesla, and I think that he had some achievement there, butnot nearly the impact I believe he potentially has. His switch (back?) to OpenAI was, in my mind, much more in keeping with where his spirit really lies. So, with that i…

How to tell if you regurgitated this comment vs being truly creative? If you can show me objectively, I’m sold.

That's not the creativity aspect, my comment is an observation, which, by definition, is a regurgitation of events.

Edit: This also demonstrates that people think (erroneously) that AI pumping out code, or content, or even essays, is inventive, but it's not.

This is merely a description and reduction, both of which AI can do, but neither of which are an invention.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#438
post #58

I would bet all of my assets of my life that AGI will not be seen in the lifetime of anyone reading this message right now. That includes anyone reading this message long after the lives of those reading it on its post date have ended. Which of course raises the interesting question of how I can make good on this bet.

genuinely curious to hear your reasoning for why this is the case. i'm always somewhere between bemused and annoyed opening the daily HN thread about AGI and seeing everyone's totally unfounded confidence in their predictions. my position is I have no idea what is going to happen.

what about the fact frontier labs are spending more compute on viral AI video slop and soon-to-be-obsoleted workplace usecases than research?

Even if you don't understand the technicals, surely you understand if any party was on the verge of AGI they wouldn't behave as these companies behave?

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#439
post #316

Earlier quoted context omitted.

Calculator can do arithmetic better than a human. Does this mean we have so called AI for half a century now?

Let's get an English major to take a calculator to the International Math Olympiad, and see how that goes.

So a sign of AGI or intelligence on par with human is the ability to solve small generic math problems? And it still requires a handler human level intellinge to be paired with, to even start solving those math problems? Is that about right?

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#440
post #391

Earlier quoted context omitted.

LLMs aren't just modeling word co-occurrences. They are recovering the underlying structure that generates word sequences. In other words, they are modeling the world. This model is quite low fidelity, but it should be very clear that they go beyond language modeling. We all know of the pelican riding a bicycle test [1]. Here's another example of how various language models view the world [2]. At this point it's just…

The "pelican on a bicycle" test has been around for six months and has been discussed a ton on the internet; that second example is fascinating but Wikipedia has infoboxes containing coordinates like 48°51′24″N 2°21′8″E (Paris, notoriously on land). How much would you bet that there isn't a CSV somewhere in the training set exactly containing this data for use in some GIS system? I think that "modeling the world" is…

>How much would you bet that there isn't a CSV somewhere in the training set exactly containing this data for use in some GIS system?

Maybe, but then I would expect more equal performance across model sizes. Besides, ingesting the data and being able to reproduce it accurately in a different modality is still an example of modeling. It's one thing to ingest a set of coordinates in a CSV indicating geographic boundaries and accurately reproduce that CSV. It's another thing to accurately indicate arbitrary points as being within the boundary or without in an entirely different context. This suggests a latent representation independent of the input tokens.

>I think that "modeling the world" is a red herring, and that fundamentally an LLM can only model its input modalities.

There are good reasons to think this isn't the case. To effectively reproduce text that is about some structure, you need a model of that structure. A strong learning algorithm should in principle learn the underlying structure represented with the input modality independent of the structure of the modality itself. There are examples of this in humans and animals, e.g. [1][2][3]

>I think a more useful definition of "model the world" is that a model needs to realize any facts that would be obvious to a person.

Seems reasonable enough, but it is at risk of being too human-centric. So much of our cognitive machinery is suited for helping us navigate and actively engage the world. But intelligence need not be dependent on the ability to engage the world. Features of the world that are obvious to us need not be obvious to an AGI that never had surviving predators or locating food in its evolutionary past. This is why I find the ARC-AGI tasks off target. They're interesting, and it will say something important about these systems when they can solve them easily. But these tasks do not represent intelligence in the sense that we care about.

>The fact that frontier models can easily be made to contradict themselves is proof enough to me that they cannot have any kind of sophisticated world model.

This proves that an LLM does not operate with a single world model. But this shouldn't be surprising. LLMs are unusual beasts in the sense that the capabilities you get largely depend on how you prompt it. There is no single entity or persona operating within the LLM. It's more of a persona-builder. What model that persona engages with is largely down to how it segmented the training data for the purposes of maximizing its ability to accurately model the various personas represented in human text. The lack of consistency is inherent to its design.

[1] https://news.wisc.edu/a-taste-of-vision-device-translates-fr...

[2] https://www.psychologicalscience.org/observer/using-sound-to...

[3] https://www.nature.com/articles/s41467-025-59342-9

Post reply on HN