Live data from Hacker News

Andrej Karpathy – It will take a decade to work through the issues with agents

dwarkesh.com

771–780 of 1001 posts

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#772

Earlier quoted context omitted.

Turing Test was a thought experiment not a real benchmark for intelligence. If you read the paper the idea originated from it is largely philosophical. As for abstract reasoning, if you look at ARC-2 it is barely capable though at least some progress has been made with the ARC-1 benchmark.

I wasn't claiming the Turing Test was a benchmark for intelligence but the ability to fool a human into thinking a machine is intelligent in conversation is still a significant milestone. I should have said "some abstract reasoning". ARC-2 looks promising.

>I wasn't claiming the Turing Test was a benchmark for intelligence but the ability to fool a human into thinking a machine is intelligent in conversation is still a significant milestone.

The Turing Test is whether it can fool a human into thinking it is talking to another human not an intelligent machine. And ironically this is becoming less true over time as people become more used to spotting the tendencies LLMs have with writing such as its frequent use of dashes or "it's not just X it is Y" type of statements.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#773
post #529
post #522

It looks like Andrej's definition of "agent" here is an entity that can replace a human employee entirely - from the first few minutes of the conversation: When you’re talking about an agent, or what the labs have in mind and maybe what I have in mind as well, you should think of it almost like an employee or an intern that you would hire to work with you. For example, you work with some employees here. When would yo…

Quite telling -- thanks for the insightful comment as always, Simon. Didn't know that, even though I've been discussing this on and off all day on Reddit. He's a smart man with well-reasoned arguments, but I think he's also a bit poisoned by working at such a huge org, with all the constraints that comes with. Like, this: You can’t just tell them something and they’ll remember it. It might take a decade to work throu…

> You can’t just tell them something and they’ll remember it.

I find it fascinating that this is the problem people consistently think we're a decade away on.

If you can't do this, you don't have employee-like AI agents, you have AI-enhanced scripting. It's basically the first thing you have to be able to do to credibly replace an actual human employee.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#774

Earlier quoted context omitted.

Why just run 20 miles then?

Because then it wouldn't be a challenge and nobody would care about the achievement.

This makes no sense.

20 miles is still a challenge, and how many people run marathons because someone else is impressed if you run 26 miles, but couldn't care less if you run 20?

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#776

This aligns with METR's Time Horizons [1], the current SOTA "Moore's Law" for AI agents: - The length of tasks AI can complete doubles every ~7 months - In 2-4 years, AIs could autonomously complete week-long projects. - In under 10 years, they might handle month-long software or knowledge work. [1] https://metr.org/blog/2025-03-19-measuring-ai-ability-to-com...

METR measures tasks, not projects. No project I've worked on had individual tasks that were supposed to take longer than 2 weeks, the PM* broke them down to sub-tasks if they were any bigger.

* At least, where we had a PM. The places I was self-directed could arguably provide an interesting comparison.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#777

Maybe I'm being too simplistic, but I think we're mixing two distinct debates. Today we have an extraordinary invention—comparable to the wheel in its time. That invention is: predictive inference over all human knowledge. Period. I don't like calling it "Artificial Intelligence" because it's not intelligence; it's a prediction system that can project responses by illuminating patterns across all human knowledge enca…

Very good point. With one caveat, though. Even though I was not there, I imagine that debates about the wheel were less heated than those we’re having about AI. I think this is because the latter is much more abstract, too close from our own consciousness etc. Wheels never challenged our place in the universe.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#779

He is an absolute treasure, I have watched all his videos more than 4 times and I don't think I would've been able to have a good mental model about deep learning without them, regardless of the amount of Bengio, Goodfellow etc lectures I have seen, none of them come even close. He is singlehandedly enabling millions of people to understand what is going on, what + and * do, actually demystifying the "wires". I just…

I agree, I think I learned the most on this topic from his videos. And before that (a while ago), it was Andrew Ng coursera's class. The latter had hands-on project, which is much better than just listening in term of retention.. I don't know if Andrej Karpathy has more structured classes somewhere.

This is a good one.

https://karpathy.ai/zero-to-hero.html

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#780

To throw two pennies in the ocean of this comment section - I’d argue we still lack schematic-level understanding of what “intelligence” even is or how it works. Not to mention how it interfaces with “consciousness”, and their likely relation to each other. Which kinda invalidates a lot of predictions/discussions of “AGI” or even in general “AI”. How can one identify Artificial Intelligence/AGI without a modicum of u…

Without going to deep into the rabbit hole, one could argue that at the first-order, intelligence is the ability to learn from experience towards a goal. In that sense, LLMs are not intelligent. They are just a (great) tool at the service of human intelligence. And so we’re just extremely far from machine intelligence.
Post reply on HN