Earlier quoted context omitted.
This world model talk is interesting, and Yann Lecunn has broached on the same topic, but the fact is there are video diffusion models that are quite good at representing the "video world" and even counterfactually and temporally coherently generating a representation of that "world" under different perturbations. In fact you can go to a SOTA LLM today, and it will do quite well at predicting the outcomes of basic co…
Photons hit a human eye and then the human came up with language to describe that and then encoded the language into the LLM. The LLM can capture some of this relationship, but the LLM is not sensing actual photons, nor experiencing actual light cone stimulation, nor generating thoughts. Its "world model" is several degrees removed from the real world. So whatever fragment of a model it gains through learning to comp…
Andrej Karpathy – It will take a decade to work through the issues with agents
811–820 of 1001 posts
Re: Andrej Karpathy – It will take a decade to work through the issues with agents
#812Earlier quoted context omitted.
You’re making the common assumption that “the algorithm“ is everything we need to get to AGI and it’s just a question of scaling.
I guess so. Is there reason to think an appropriate algorithm and scale can't do that?
Re: Andrej Karpathy – It will take a decade to work through the issues with agents
#813Earlier quoted context omitted.
That's a model. Not a higher-order model like most humans use, but it's still a model.
Yes, not of the world, but of the ingested text. Almost verbatim what I wrote.
Are you all so terminally nerd brained you can’t see the obvious
Re: Andrej Karpathy – It will take a decade to work through the issues with agents
#814Why does everyone have such short timelines to show progress? So what if it takes 50 years to develop, we’ll have AGI for the next million years
I don't know how much wish fulfilment there is in people's timelines.
Re: Andrej Karpathy – It will take a decade to work through the issues with agents
#815Earlier quoted context omitted.
I’m with you on the energy and limitations, and even on the moving of goalposts. I’d like to add that I think limit definition of AGI has jumped the shark though and is already at ASI, since we expect our machine to exhibit professional level acumen across such a wide range of knowledge that it would be similar to the 0.01 percent top career scholars and engineers, or even above any known human capacity just due to b…
As much regulatory measures as possible seems good. This things are not toys.
Re: Andrej Karpathy – It will take a decade to work through the issues with agents
#816Earlier quoted context omitted.
As much regulatory measures as possible seems good. This things are not toys.
Yeah, it’s definitely some kind of new chapter. It’s reducing hiring, and will drive unemployment, no matter what people are saying. It’s a poison pill in a way, since no one will hire junior staff anymore. The reliance on AI will skyrocket as experienced staff ages out and there are no replacements coming up through the ranks.
Re: Andrej Karpathy – It will take a decade to work through the issues with agents
#817Earlier quoted context omitted.
This world model talk is interesting, and Yann Lecunn has broached on the same topic, but the fact is there are video diffusion models that are quite good at representing the "video world" and even counterfactually and temporally coherently generating a representation of that "world" under different perturbations. In fact you can go to a SOTA LLM today, and it will do quite well at predicting the outcomes of basic co…
Photons hit a human eye and then the human came up with language to describe that and then encoded the language into the LLM. The LLM can capture some of this relationship, but the LLM is not sensing actual photons, nor experiencing actual light cone stimulation, nor generating thoughts. Its "world model" is several degrees removed from the real world. So whatever fragment of a model it gains through learning to comp…
Neither is animal brain. It's processing the signals produced by the sensors. Once the world model is programmed/auto-built in the brain, it doesn't matter if it's sensing real photons, it just has input pins like a transistor or arguments of a function. As long as we provide the arguments, it doesn't matter how those arguments are produced. LLMs are not different in that aspect.
> nor generating thoughts
They do during the chain-of-thought process. Generally there's no incentive to let an LLM keep mulling over a topic as that is not useful to the humans and they make money only when their gears start turning in response to a question sent by a human. But that doesn't mean that LLM doesn't have capability to do that.
> Its "world model" is several degrees removed from the real world.
Just because animal brain has tools called sensors that it can get data from world without external stimuli, it doesn't mean that it's any closer to the world than an LLM. It's still getting ultra processed signals to feed to its own programming. Similarly, LLMs do interact with real world through tools as agent.
> So whatever fragment of a model it gains through learning to compress that causal chain of events does not mean much when it cannot generate the actual causal chain.
Again, a person who has gone blind, still has the world model created by the sight. This person can also no longer generate the chain of events that led to creation of that sight model. It still doesn't mean that this person's world model has become inferior.
Re: Andrej Karpathy – It will take a decade to work through the issues with agents
#818Earlier quoted context omitted.
This world model talk is interesting, and Yann Lecunn has broached on the same topic, but the fact is there are video diffusion models that are quite good at representing the "video world" and even counterfactually and temporally coherently generating a representation of that "world" under different perturbations. In fact you can go to a SOTA LLM today, and it will do quite well at predicting the outcomes of basic co…
Photons hit a human eye and then the human came up with language to describe that and then encoded the language into the LLM. The LLM can capture some of this relationship, but the LLM is not sensing actual photons, nor experiencing actual light cone stimulation, nor generating thoughts. Its "world model" is several degrees removed from the real world. So whatever fragment of a model it gains through learning to comp…
I'll give you the brain is currently better at the world modelling stuff but Genie 3 is pretty impressive.
Re: Andrej Karpathy – It will take a decade to work through the issues with agents
#819Earlier quoted context omitted.
In my view 'understand' is a folk psychology term that does not have a technical meaning. Like 'intelligent', 'beautiful', and 'interesting'. It usefully labels a basket of behaviors we see in others, and that is all it does. In this view, if a machine performs a task as well as a human, it understands it exactly as much as a human. There's no problem of how to do understanding, only how to do tasks. The 'problem' me…
> if a machine performs a task as well as a human, it understands it exactly as much as a human. I think you're right, except that the ones judging "as well as a human" are in fact humans, and humans have expectations that expand beyond the specs. From the narrow perspective of engineering specifications or profit generated, a robot/AI may very well be exactly as understanding as a human. For the people which interac…
It is like saying the airplane understands how to fly.
"You disagree? Well lets see you fly! You are saying the airplane doesn't understand how to fly and you can't even fly yourself?"
This would be confusing the fact humans built the flying machine and the flying machine doesn't understand anything.
Re: Andrej Karpathy – It will take a decade to work through the issues with agents
#820Earlier quoted context omitted.
Is that not insane?
Both Karpathy and must have explained it many times. The additional sensors add more signal than noise in the end. You also then have to decide which sensor system is correct any time they disagree. Also the entire road system is designed for vision. Lidar cannot read signs, see colors, etc. Humans can drive with two eyes. It's not insane to think computers can do it with 7 or 8 cameras. As someone who has used Tesla…