Live data from Hacker News

Andrej Karpathy – It will take a decade to work through the issues with agents

dwarkesh.com

291–300 of 1001 posts

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#291
post #73

Earlier quoted context omitted.

like leveling to 99 in old school runescape

Or Diablo 2

i don't remember the end-game of the original Diablo; however, in diablo III and IV everyone i've tried to play the game gets bored in the run up to max level. I always tell them "i skip that part as much as possible, because that's not the game. That's just the story!"

Once you hit max level in III and IV, the game actually "begins."

and to explain the Diablo 2 Reference, the amount of time/effort it takes to go from level 98 to level 99 (the max level), is the same amount of time it takes to go from level 1 to level 98. I've heard "2 weeks" as a rough estimate of "unhealthy playtime", at least solo.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#292
post #279

Earlier quoted context omitted.

The interview which I've watched recently with Rich Sutton left me with the impression that AGI is not just a matter of adding more 9s. The interviewer had an idea that he took for granted: that to understand language you have to have a model of the world. LLMs seem to udnerstand language therefore they've trained a model of the world. Sutton rejected the premise immediately. He might be right in being skeptical here…

I think this a useful challenge to our normal way of thinking. At the same time, "the world" exists only in our imagination (per our brain). Therefore, if LLMs need a model of a world, and they're trained on the corpus of human knowledge (which passed through our brains), then what's the difference, especially when LLMs are going back into our brains anyway?

Language isn't thought. It's a representation of thought.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#293
post #246
post #226

Earlier quoted context omitted.

There is some evidence from Anthropic that LLMs do model the world. This paper[0] tracing their "thought" is fascinating. Basically an LLM translating across languages will "light up" (to use a rough fMRI equivalent) for the same concepts (e.g. bigness) across languages. It does have clusters of parameters that correlate with concepts, not just randomly "after X word tends to have Y word." Otherwise you would expect…

Right, but modeling the structure of language is a question of modeling word order and binding affinities. It's the Chinese Room thought experiment - can you get away with a form of "understanding" which is fundamentally incomplete but still produces reasonable outputs? Language in itself attempts to model the world and the processes by which it changes. Knowing which parts-of-speech about sunrises appear together an…

> Knowing which parts-of-speech about sunrises appear together and where is not the same as understanding a sunrise

What does "understanding a sunrise" mean though? Arguments like this end up resting on semantics or tautology, 100% of the time. Arguments of the form "what AI is really doing" likewise fail because we don't know what real brains are "really" doing either.

I mean, if we knew how to model human language/reasoning/whatever we'd just do that. We don't, and we can't. The AI boosters are betting that whatever it is (that we don't understand!) is an emergent property of enough compute power and that all we need to do is keep cranking the data center construction engine. The AI pessimists, you among them, are mostly just arguing from ludditism: "this can't possibly work because I don't understand how it can".

Who the hell knows, basically. We're at an interesting moment where technology and the theory behind it are hitting the wall at the same time. That's really rare[1], generally you know how something works and applying it just a question of figuring out how to build a machine.

[1] Another example might be some of the chemistry fumbling going on at the start of the industrial revolution. We knew how to smelt and cast metals at crazy scales well before we knew what was actually happening. Stuff like that.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#295
post #261

I always get a weird feeling when AI researchers and CS people start talking about comparisons between human brains and AI/computers Why is there a presumption that we (as people who have only studied CS) know enough about biology/neuroscience/evolution to make these comparisons/parallels/analogies? I enjoy the discussions but I always get the thought in the back of my head "...remember you're listening to 2 CS major…

From personal experience making the same comparisons during undergrad, I think it just comes down to the availability of conceptual models. If the brain does X, there's a good chance that a computer does something that looks like X, or that X could be recreated through steps Y & Z, etc. Once I started to realize just how much of the brain is inscrutable, because it is a machine operating on chemicals instead of stric…

Lucky for all of us we're alive during a "quantum" thing! Which has been an idea since at least the mid 1990s as i first saw it in a 2600 around that time...

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#296
post #58

I would bet all of my assets of my life that AGI will not be seen in the lifetime of anyone reading this message right now. That includes anyone reading this message long after the lives of those reading it on its post date have ended. Which of course raises the interesting question of how I can make good on this bet.

Should probably just short nvidia

"just short nvidia" is not simple. Even if you believe it is overvalued, and you are correct, a short is a specific bet that the market will realize that fact in a precise amount of time. There are very significant risks in short selling, and famously, the market can stay irrational longer than you can remain solvent.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#297
post #289

Earlier quoted context omitted.

>Why is there a presumption that we (as people who have only studied CS) know enough about biology/neuroscience/evolution to make these comparisons? Hubris.

The hubris here isn't CS people making comparisons, it's assuming biological substrate matters. Your brain is doing computation with neurotransmitters instead of transistors. So what? The "chemicals not electricity" distinction is pure carbon chauvinism, like insisting hydraulic computers can't be compared to electronic ones because water isn't electricity. Evolution didn't discover some mystical process that imbues…

Case in point.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#298
post #201

Not a decade. More like a century, and that is if society figures itself out enough to do some engineering on a planetary scale, and quantum computing is viable. Fundamentally, AGI requires 2 things. First it needs to be able to operate without information, learning as it goes. The core kernel should be such that it doesn't have any sort of training on real world concepts, only general language parsing that it can us…

You want a "core kernel" with "general language parsing" but no training on real-world concepts. Read that sentence again. Slowly. What do you think "general language parsing" IS if not learned patterns from real-world data? You're literally describing a transformer and then saying we need to invent it. And your TLS example is deranged. You want an agent to discover the TLS protocol by randomly sending ethernet packe…

>What do you think "general language parsing" IS if not learned patterns from real-world data?

I want you to hertograize the enpostule by brasetting the leekerists, while making sure that the croalbastes are not exhibiting any ecrocrafic effects

Whatever you understand about that task, is what a kernel will "understand" as well. And however you go about solving it, the kernel will also will follow similar patterns of behaviour (starting with figuring out what hertrograize means, which then leads to other tasks, and so on)

>You want an agent to discover the TLS protocol by randomly sending ethernet packets? The combinatorial search space is so large this wouldn't happen before the sun explodes.

In pure combination, yes. In smart directed intelligent search, no. Ideally the kernel could listen for incoming traffic, and figure out patterns based on that. But the point is that the kernel should figure out that listening for traffic is optimal without you specifically telling it, because it "understands" the concept of other "entities" communicating with it and that communication is bound to be in a structured format, and has internal reward systems in place for figuring it out through listening rather than expending energy brute force searching.

Whatever that process is, it will get applied to much harder problems identically.

>Transformers already ARE general algorithms with zero hardcoded linguistic knowledge. The architecture doesn't know what a noun is. It doesn't know what English is. It learns everything from data through gradient descent. That's the entire damn point.

It doesn't learn what a noun is or english is, its a statistical mapping that just tends to work well. LLMs are just efficient look up maps. Look up maps can go only so far as to interpolate on the knowledge encoded within them. These can simulate intelligence in the sense of recursive lookups, but fundamentally that process is very guided, hence all the manual things like prompt engineering, mcp servers, agents, skills and so on.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#299
post #68

>What takes the long amount of time and the way to think about it is that it’s a march of nines. Every single nine is a constant amount of work. Every single nine is the same amount of work. When you get a demo and something works 90% of the time, that’s just the first nine. Then you need the second nine, a third nine, a fourth nine, a fifth nine. While I was at Tesla for five years or so, we went through maybe three…

I also quite like the way he puts it. However, from a certain point onward, the AI itself will contribute to the development—adding nines—and that’s the key difference between this analogy of nines in other systems (including earlier domain‑specific ML ones) and the path to AGI. That's why we can expect fast acceleration to take off within two years.

I doubt this. General intelligence will be a step change not a gentle ramp. If we get to an architecture intelligent enough to meaningfully contribute to AI development, we'll have already made it. It'll simply be a matter of scale. There's no 99% AGI that can help build 100% AGI but for some reason can't drive a car or cook a meal or work an office job.

Re: Andrej Karpathy – It will take a decade to work through the issues with agents

#300

Earlier quoted context omitted.

Reminds me of a time-honored aphorism in running: A marathon consists of two halves: the first 20 miles, and then the last 10k (6.2mi) when you're more sore and tired than you've ever been in your life.

Why just run 20 miles then?

because that'd be quitting the race with 6.2 miles left to go
Post reply on HN