Earlier quoted context omitted.
> Basically an LLM translating across languages will "light up" for the same concepts across languages Which is exactly what they are trained to do. Translation models wouldn't be functional if they are unable to correlate an input to specific outputs. That some hiddel-layer neurons fire for the same concept shouldn't come as a surprise, and is a basic feature required for the core functionality.
And if it is true that the language is just the last step after the answer is already conceptualized, why do models perform differently in different languages? If it was just a matter of language, they’d have the same answer but just with a broken grammar, no?
Andrej Karpathy – It will take a decade to work through the issues with agents
581–590 of 1001 posts
Re: Andrej Karpathy – It will take a decade to work through the issues with agents
#582Re: Andrej Karpathy – It will take a decade to work through the issues with agents
#583Earlier quoted context omitted.
I agree with this. A metaphor I like is that the reason why humans say the night sky is beautiful is because they see that it is, whereas an LLM says it because it’s been said enough times in its training data.
I mean, I think the reason I would say the night sky is “beautiful” is because the meaning of the word for me is constructed from the experiences I’ve had in which I’ve heard other people use the word. So I’d agree that the night sky is “beautiful”, but not because I somehow have access to a deeper meaning of the word or the sky than an LLM does. As someone who (long ago) studied philosophy of mind and (Chomskian) li…
Ok but you don’t look at every night sky or every sunset and say “wow that’s beautiful”
There’s a quality to it - not because you heard someone say it but because you experience it
Re: Andrej Karpathy – It will take a decade to work through the issues with agents
#584Re: Andrej Karpathy – It will take a decade to work through the issues with agents
#585Re: Andrej Karpathy – It will take a decade to work through the issues with agents
#586Re: Andrej Karpathy – It will take a decade to work through the issues with agents
#587With all due respect, what does it say about us that „famous researcher voices his speculative opinion“ is an instant top 1 on hackernews?
Re: Andrej Karpathy – It will take a decade to work through the issues with agents
#588>What takes the long amount of time and the way to think about it is that it’s a march of nines. Every single nine is a constant amount of work. Every single nine is the same amount of work. When you get a demo and something works 90% of the time, that’s just the first nine. Then you need the second nine, a third nine, a fourth nine, a fifth nine. While I was at Tesla for five years or so, we went through maybe three…
The interview which I've watched recently with Rich Sutton left me with the impression that AGI is not just a matter of adding more 9s. The interviewer had an idea that he took for granted: that to understand language you have to have a model of the world. LLMs seem to udnerstand language therefore they've trained a model of the world. Sutton rejected the premise immediately. He might be right in being skeptical here…
Re: Andrej Karpathy – It will take a decade to work through the issues with agents
#589Re: Andrej Karpathy – It will take a decade to work through the issues with agents
#590Earlier quoted context omitted.
I agree with this. A metaphor I like is that the reason why humans say the night sky is beautiful is because they see that it is, whereas an LLM says it because it’s been said enough times in its training data.
I mean, I think the reason I would say the night sky is “beautiful” is because the meaning of the word for me is constructed from the experiences I’ve had in which I’ve heard other people use the word. So I’d agree that the night sky is “beautiful”, but not because I somehow have access to a deeper meaning of the word or the sky than an LLM does. As someone who (long ago) studied philosophy of mind and (Chomskian) li…
People are just really really complex machines.
However there are clearly qualitative differences between the human mind and any machines we know of yet, and those qualitative differences are emergent properties, in the same way that a rabbit is qualitatively different than a stone or a chunk of wood.
I also think most of the recent AI experts/optimists underestimate how complex the mind is. I'm not at the cutting edge of how LLMs are being trained and architected, but the sense I have is we haven't modelled the diversity of connections in the mind or diversity of cell types. E.g. Transcriptomic diversity of cell types across the adult human brain (Siletti et al., 2023, Science)