It was provocative back then, and still is!
The idea of having my consciousness continue its existence indefinitely in somebody's kubernetes cluster, however, seems like a very special vision of hell.
21–30 of 186 posts
It was provocative back then, and still is!
The idea of having my consciousness continue its existence indefinitely in somebody's kubernetes cluster, however, seems like a very special vision of hell.
Aren't we there yet? George Hotz has a nice series of posts on this: https://geohot.github.io/blog/jekyll/update/2020/08/20/a-sil... https://geohot.github.io/blog/jekyll/update/2022/02/17/brain... https://geohot.github.io/blog/jekyll/update/2023/04/26/a-per... I sorted them by date.
Moravac (in the linked paper): "In both cases, the evidence for an intelligent mind lies in the machine's performance, not its makeup." Do you agree? I'm much less keen to ascribe "intelligence" to large, pretrained language models given that I know how primitive their training regime is compared to a scenario where I might have been "blended" by their ability to "chat" (double quote here since I know ChatGPT and the…
Given their performance, I think it is important to pay attention to their weirdnesses — I can call them "intelligent" or "dumb" without contradiction depending on which specific point is under consideration.
Transistors outpace biological synapses by the same degree to which a marathon runner outpaces continental drift. This speed difference is what allows computers to read the entire text content of the internet on a regular basis, whereas a human can't read all of just the current version of the English language Wikipedia once in their lifetime.
But current AI is very sample-inefficient: if a human were to read as much as an LLM, they would be world experts at everything, not varying between "secondary school" and "fresh graduate" depending on the subject… but even that description is misleading, because humans have the System 1/System 2[0] distinction and limited attention[1], whereas LLMs pay attention to approximately everything in the context window and (seem to) be at a standard between our System 1 and System 2.
If you're asking about an LLM's intelligence because you want to replace an intern, then the AI are intelligent; but if you're asking because you want to know how many examples they need in order to decode North Sentinelese or Linear A, then (from what I understand) these AI are extremely stupid.
It doesn't matter if a submarine swims[2], it still isn't going to fit into a flooded cave.
[0] https://en.wikipedia.org/wiki/Thinking,_Fast_and_Slow
[1] https://youtu.be/vJG698U2Mvo?si=omf3xleqPw5u6Y2k
https://youtu.be/ubNF9QNEQLA?si=Ja-9Ak4iCbcxbWdh
https://youtu.be/v3iPrBrGSJM?si=9cKHXEvEGl764Efa
https://www.americanbar.org/groups/intellectual_property_law...
[2] https://www.goodreads.com/quotes/32629-the-question-of-wheth...
Aren't we there yet? George Hotz has a nice series of posts on this: https://geohot.github.io/blog/jekyll/update/2020/08/20/a-sil... https://geohot.github.io/blog/jekyll/update/2022/02/17/brain... https://geohot.github.io/blog/jekyll/update/2023/04/26/a-per... I sorted them by date.
> We’ll use the estimates of 100 billion neurons and 100 trillion synapses for this post. That’s 100 teraweights.
... or maybe actual synapses cannot be described with a single weight.
> The max firing rate seems to be 200hz [4]. I really want an estimate of “neuron lag” here, but let’s upper bound it at 5ms.
... but ANNs output float activations per pass. Biological neurons encode values with sequences of spikes which vary in timing, so the firing rate doesn't on its own tell you the rate at which neurons can communicate new values.
> Remember also, that the brain is always learning, so it needs to be doing forward and backward passes. I’m not exactly sure why they are different, but [6] and [7] point to the backward pass taking 2x more compute than the forward pass.
... but the brain probably isn't doing backprop, in part because it doesn't get to observe the 'correct' output, compute a loss, etc, and because the brain isn't a DAG.
Moravac (in the linked paper): "In both cases, the evidence for an intelligent mind lies in the machine's performance, not its makeup." Do you agree? I'm much less keen to ascribe "intelligence" to large, pretrained language models given that I know how primitive their training regime is compared to a scenario where I might have been "blended" by their ability to "chat" (double quote here since I know ChatGPT and the…
I agree with Moravec. As he points out a bit later on: > Only on the outside, where they can be appreciated as a whole, will the impression of intelligence emerge. A human brain, too, does not exhibit the intelligence under a neurobiologist's microscope that it does participating in a lively conversation. We only have fuzzy definitions of "intelligence", not any essential, unambiguous things we can point to at a minu…
(Chollet, 2019, https://arxiv.org/pdf/1911.01547.pdf)
Priors here means how targeted is the model design to the task. Experience means how large is the necessary training set. Generalization difficulty is how hard is the task.
So intelligence is defined as ability to learn a large number of tasks with as little experience and model selection as possible. If it's a skill only possible because your model already follows the structure of the problem, then it won't generalize. If it requires too much training data, it's not very intelligent. If it's just a set number of skills and can't learn new ones quickly, it's not intelligent.
Aren't we there yet? George Hotz has a nice series of posts on this: https://geohot.github.io/blog/jekyll/update/2020/08/20/a-sil... https://geohot.github.io/blog/jekyll/update/2022/02/17/brain... https://geohot.github.io/blog/jekyll/update/2023/04/26/a-per... I sorted them by date.
Hmm, this reasoning is making a lot of really questionable assumptions: > We’ll use the estimates of 100 billion neurons and 100 trillion synapses for this post. That’s 100 teraweights. ... or maybe actual synapses cannot be described with a single weight. > The max firing rate seems to be 200hz [4]. I really want an estimate of “neuron lag” here, but let’s upper bound it at 5ms. ... but ANNs output float activations…
Moravec also wrote a book much along these lines in the late 80's (Mind Children). If I recall correctly, a good part of it was also about the idea of "transferring" human consciousness into a machine "host". The idea being that a sufficiently advanced computer would be able to somehow make sense of a human's neuron mappings and then "continue running" as that individual. It was provocative back then, and still is! T…
Moravec also wrote a book much along these lines in the late 80's (Mind Children). If I recall correctly, a good part of it was also about the idea of "transferring" human consciousness into a machine "host". The idea being that a sufficiently advanced computer would be able to somehow make sense of a human's neuron mappings and then "continue running" as that individual. It was provocative back then, and still is! T…
Also a provocative read.
Moravac (in the linked paper): "In both cases, the evidence for an intelligent mind lies in the machine's performance, not its makeup." Do you agree? I'm much less keen to ascribe "intelligence" to large, pretrained language models given that I know how primitive their training regime is compared to a scenario where I might have been "blended" by their ability to "chat" (double quote here since I know ChatGPT and the…
Reducing the capability of the human brain to performance alone is too simplistic, especially when looking at LLM's. Even if we would assign some intelligence to LLM's, they need a 400w GPU at inference time, and several orders of magnitude more of those at training time. The human brain runs constanly at ~20w. I highly doubt you'd be able to get even close to that kind of performance with current manufacturing proce…
At a low level: We take an analog component, then drive it in a way that lets us treat it as digital, then combine loads of them together so we can synthesise a low-resolution approximation of an analog process.
At a higher level: We don't really understand how our brains are architected yet, just that it can make better guesses from fewer examples than our AI.
Also, 400 W of electricity is generally cheaper than 20 W of calories (let alone the 38-100 W rest of body needed to keep the brain alive depending on how much of a couch potato the human is).
Moravac (in the linked paper): "In both cases, the evidence for an intelligent mind lies in the machine's performance, not its makeup." Do you agree? I'm much less keen to ascribe "intelligence" to large, pretrained language models given that I know how primitive their training regime is compared to a scenario where I might have been "blended" by their ability to "chat" (double quote here since I know ChatGPT and the…
This is reminding me again of The Bitter Lesson. http://www.incompleteideas.net/IncIdeas/BitterLesson.html
Foundational changes are of course harder, but it does not mean we should drop it all together.