Live data from Hacker News

Andrej Karpathy – It will take a decade to work through the issues with agents

dwarkesh.com

991–1000 of 1001 posts

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#991

Earlier quoted context omitted.

The interview which I've watched recently with Rich Sutton left me with the impression that AGI is not just a matter of adding more 9s. The interviewer had an idea that he took for granted: that to understand language you have to have a model of the world. LLMs seem to udnerstand language therefore they've trained a model of the world. Sutton rejected the premise immediately. He might be right in being skeptical here…

This world model talk is interesting, and Yann Lecunn has broached on the same topic, but the fact is there are video diffusion models that are quite good at representing the "video world" and even counterfactually and temporally coherently generating a representation of that "world" under different perturbations. In fact you can go to a SOTA LLM today, and it will do quite well at predicting the outcomes of basic co…

> Animal brains such as our own have evolved to compress information about our world to aide in survival.

Key question is what are the "selection pressures" that drive the "evolution" of LLMs? In the case of robotics, there's a "survival of task completion" which usually has some physical goal, like assembling a part correctly or scoring a goal on a soccer field. One of the selection pressures driving LLM evolution is that the dual of always answering with something AND continuing the conversation (engagement). You can imagine how those two selection pressures yield outcomes that don't represent the world in a "real" sense.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#992
post #967

Earlier quoted context omitted.

Just, ftr, endorphins cannot pass the blood brain barrier http://hopkinsmedicine.org/health/wellness-and-prevention/th...

So a runner’s high is more like a literal high then? Interesting

Ha! Endorphins are "endogenous opioid peptides produced by the pituitary and hypothalamus glands that function as the body's natural painkillers and mood regulators".

"They are part of the endogenous opioid system", so either way was talking about literal highs.

The endocannabinoid system (I hope that I have the spelling correct) is a relatively recent discovery (1980s on), and is quite fascinating on how integral to the human body it is

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#993
post #803

Earlier quoted context omitted.

Oh look, people with skin in the AI game insist AI is not a massive bubble. More news at 11.

We’re a regular old SaaS company that has figured out how to add massive value using AI. I am making no statements about valuations and bubbles. I’m actually guessing there is some bubble / overhype. That doesn’t mean it isn’t still incredibly valuable.

[dead]

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#994

Earlier quoted context omitted.

As someone who is invested in researching said math, I can say with some confidence that it does not exist, or at least not in the form claimed here. That's the whole problem. I would be ecstatic if it did though, so if anyone has any examples or rebuttal, I would very much appreciate it.

Let me clarify. I was too vague and definitely did not express things accurately. That is on me. We have the math to show that it can be impossible to distinguish two explanations through data processing alone. We have examples of this in science, a long history of it in fact. Fundamentally there is so much that we cannot conclude from processing data alone. Science (the search of knowledge) is active. It doesn't req…

My apologies for the much delayed reply as I have recently found myself with little extra time to post adequate responses. Your critiques are very interesting to ponder, so I thank you for posting them. I did want to respond to this one though.

I believe all of my counterarguments center around my current viewpoint that given the rapid rate of progress involved on the engineering side, it is no longer reasonable in deep learning theory to consider what is possible, and it is more interesting to try to outline hard limitations. This emposes a stark contrast between deep learning and classical statistics, as the boundaries in the latter are very clear and are not shared by the former.

I want to stress that at present, nearly every conjectured limitation of deep learning over the last several decades has fallen. This includes many back of the napkin, "clearly obvious" arguments, so I'm wary of them now. I think the skepticism all along has been fueled in response to hype cycles, so we must be careful not to make the same mistakes. There is far too much empirical evidence available to counter precise arguments against the claim that there is an underlying understanding within these models, so it seems we must resort to the imprecise to continue the debate.

Scaling, along one axis, suggests a high polynomial degree of additional compute (not exponential) is required for increasing improvements, this is true. But the progress over the last few years has occurred due to the discovery of new axes to scale on, which further reduces the error rate and improves performance. There are still many potential axes left untapped. What is significant about scaling to me is not how much additional compute is required, but the fact that the predicted bottom at the moment is very, very low, far lower than anything else we have ever seen, and that doesn't require any more data than we currently have. That should be cause for concern until we find a better lower bound.

> We all know that these new LLMs aren't dramatic improvements off their previous versions

No, I don't agree. This may be evident to many, but to some, the differences are stark. Our perceived metrics of performance are nonlinear and person-dependent, and these major differences can be imperceptible to most. The vast majority of attempts at providing more regular metrics or benchmarks that are not already saturated have shown that LLM development is not slowing down by any stretch. I'm not saying that LLMs will "go to the moon". But I don't have anything concrete to say they cannot either.

> We have the math to show that it can be impossible to distinguish two explanations through data processing alone.

Actually, this is a really great point, but I think this highlights the limitations of benchmarks and the requirements of capacity-based, compression-based, or other types of alternative data-independent metrics. With these in tow, it can be possible to distinguish two explanations. This could be a fruitful line of inquiry.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#995
post #288

Earlier quoted context omitted.

I think a ton of people see a line going up and they think exponential. When in Reality, the vast majority of the time it’s actually logistic.

Given the physical limits of the universe and our planet in particular, yeah, this is pretty much always true. The interesting question is: what is that limit, and: how many orders of magnitude are we away from leveling off?

For reasoning the data suggests models are in the logistic domain?

Grok 3 - Grok 3 reasoning: 15% increase in training compute for a 25% uplift in inteligence

Grok 3 reasoning - Grok 4: 80% increase in training compute for a 15% uplift in inteligence.

Inteligence: Source Atrifical Analysis

Training Compute: Source https://youtu.be/MtYsUdfZPMA?t=162

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#996

This aligns with METR's Time Horizons [1], the current SOTA "Moore's Law" for AI agents: - The length of tasks AI can complete doubles every ~7 months - In 2-4 years, AIs could autonomously complete week-long projects. - In under 10 years, they might handle month-long software or knowledge work. [1] https://metr.org/blog/2025-03-19-measuring-ai-ability-to-com...

METR uses a 50% success rate in that analysis, beecause the models are non-determistic.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#997
post #784

Unless someone can show me some sort of "Moore's law" for LLM's, saying it will "take a decade" sounds more to me like it could "take 10 years for the next 20 years".

METR kinda has been described as a Moore's Law for LLMs, but personally I think the financial environment around AI will break within 2 years — which still represents a huge increase in capabilities, but isn't a decade. • Text and graphs: https://metr.org/blog/2025-03-19-measuring-ai-ability-to-com... • Video interview: https://www.youtube.com/watch?app=desktop&v=evSFeqTZdqs That said, I've not seen work that looks p…

METR's study uses a 50% sucess rate. For most enterpise applications a 50% pass rate is unacceptable in automation, more like >99.9%.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#998

I want to be heretical and say that Karpathy hasn't worked in a frontier lab since 2020 and missed all the greatness of the last years. Humans are humans are humans.

Matt Levine never worked in crypto, but he knew SBF/FTX was a scam. Humans are humans.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#999

Earlier quoted context omitted.

I’m not saying he’s saying agents aren’t useful at all. It’s literally in the quotes I provided that he says they are useful for some subset of tasks. I’m saying that he is answering the question “are agents useful at all”. not “can agents replace humans”. His answer is mostly not. He generally prefers autocomplete. But they are useful for some limited tasks.

Saying "He’s talking about whether agents are currently useful at all" is negatively loaded. It is very easy to take that and assume the answer is "no" based on the "at all". If you wanted to be more neutral, you could have said something like "He's also questioning how useful agents really are today". That wouldn't have implied that they're not useful at all, but instead that they're less useful than people are clai…

That question doesn’t do enough to highlight how wrong the OP’s interpretation was. He’s going far beyond just stating that agents are less useful than people are claiming. Less useful than people are claiming fits the OP’s interpretation.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#1000
post #940

Earlier quoted context omitted.

Long form engineering tasks aren’t doable yet without supervision. But I can say in our shop, we won’t be hiring any more junior devs, ever, except as (in my region, free) interns or because of some extraordinary capabilities, insights, or skills. There just isn’t any business case for hiring junior devs to do the grunt work anymore. But, the vast majority of work that is done in the world is not in the same order of…

What does junior or senior have anything to do with it ? I would think a smarter junior will run circles around a dumber senior engineer with LLM autocomplete.

If you’re hiring dumb senior engineers you’re holding it wrong lol. Using LLMs is a lot like delegating to a team from a skills perspective, so it favors extensive domain knowledge. You don’t just commit whatever it writes, just like you wouldn’t commit what a junior dev writes without scrutiny. Experience makes that scrutiny more valuable and effective.
Post reply on HN