Live data from Hacker News

A non-anthropomorphized view of LLMs

addxorrol.blogspot.com

391–400 of 432 posts

Re: A non-anthropomorphized view of LLMs

#391
post #154
post #126

Earlier quoted context omitted.

I kinda agree with both of you. It might be a required abstraction, but it's a leaky one. Long before LLMs, I would talk about classes / functions / modules like "it then does this, decides the epsilon is too low, chops it up and adds it to the list". The difference I guess it was only to a technical crowd and nobody would mistake this for anything it wasn't. Everybody know that "it" didn't "decide" anything. With AI…

Agreeing with you, this is a "can a submarine swim" problem IMO. We need a new word for what LLMs are doing. Calling it "thinking" is stretching the word to breaking point, but "selecting the next word based on a complex statistical model" doesn't begin to capture what they're capable of. Maybe it's cog-nition (emphasis on the cog).

This is a total non-problem that has been invented by people so they have something new and exciting to be pedantic about.

When we need to speak precisely about a model and how it works, we have a formal language (mathematics) which allows us to be absolutely specific. When we need to empirically observe how the model behaves, we have a completely precise method of doing this (running an eval).

Any other time, we use language in a purposefully intuitive and imprecise way, and that is a deliberate tradeoff which sacrifices precision for expressiveness.

Re: A non-anthropomorphized view of LLMs

#392

Earlier quoted context omitted.

> It wasn't a simple brute force. I think you misunderstood me. "Simple" is the key word here, right? You agree that it is still under the broad class of "brute force"? I'm not saying Claude is naively brute forcing. In fact, with lack of interpretibility of these machines it is difficult to say what kind of optimization it is doing and how complex that it (this was a key part tbh). My point was to help with this > I…

I have no idea why some people take so much offense to rhe fact humans are just another machine, there's no reason why another machine can't surpass it here as in all other aveneus machines have already. Many of the reasons people give for llms not being conscious are just as applicable to humans too.

Absolutely possible (I’d say even likely) for humans to be surpassed by machines who have better recall and storage already.

I’m highly skeptical this will happen with llms though, their output is superficially convincing but without depth and creativity.

Re: A non-anthropomorphized view of LLMs

#393

Earlier quoted context omitted.

Author of the original article here. What hidden state are you referring to? For most LLMs the context is the state, and there is no "hidden" state. Could you explain what you mean? (Apologies if I can't see it directly)

You wrote this article and you're not familiar with hidden states?

I am not aware that an LLM contains any.

Re: A non-anthropomorphized view of LLMs

#394

Earlier quoted context omitted.

The "point" of not anthropomorphizing is to refrain from judgement until a more solid abstraction appears. The problem with explaining LLMs in terms of human behaviour is that, while we don't clearly understand what the LLM is doing, we understand human cognition even less! There is literally no predictive power in the abstraction "The LLM is thinking like I am thinking". It gives you no mechanism to evaluate what ta…

> It is like someone inventing the aeroplane and someone looks at it and says "oh, it's flying, I guess it's a bird". It's not a bird! We tried to mimic birds at first; it turns out birds were way too high-tech, and too optimized. We figured out how to fly when we ditched the biological distraction and focused on flight itself. But fast forward until today, we're reaching the level of technology that allows us to bui…

A machine that emulates a bird is indeed a mechanical bird. We can say what emulating a bird is because we know, at least for the purpose of flying, what a bird is and how it works. We (me, you, everyone else) have no idea how thinking works. We do not know what consciousness is and how it operates. We may never know. It is deranged gibberish to look at an LLM and say "well, it does some things I can do some of the time, so I suppose it's a digital mind!". You have to understand the thing before you can say you're emulating it.

Re: A non-anthropomorphized view of LLMs

#395

Earlier quoted context omitted.

> It wasn't a simple brute force. I think you misunderstood me. "Simple" is the key word here, right? You agree that it is still under the broad class of "brute force"? I'm not saying Claude is naively brute forcing. In fact, with lack of interpretibility of these machines it is difficult to say what kind of optimization it is doing and how complex that it (this was a key part tbh). My point was to help with this > I…

I have no idea why some people take so much offense to rhe fact humans are just another machine, there's no reason why another machine can't surpass it here as in all other aveneus machines have already. Many of the reasons people give for llms not being conscious are just as applicable to humans too.

I don't think the question is if humans are a machine or not but rather what is meant by machine. Most people interpret it as meaning deterministic and thus having no free will. That's probably not what you're trying to convey so might not be the best word to use.

But the question is what is special about the human machine? What is special about the animal machine? These are different from all the machines we have built. Is it complexity? Is it indeterministic? Is it more? Certainly these machines have feelings, and we need to account for them when interacting with them.

Though we're getting well off topic from determining if a duck is a duck or is a machine (you know what I mean by this word and that I don't mean a normal duck)

Re: A non-anthropomorphized view of LLMs

#396

Earlier quoted context omitted.

"Not conscious" is a silly claim. We have no agreed-upon definition of "consciousness", no accepted understanding of what gives rise to "consciousness", no way to measure or compare "consciousness", and no test we could administer to either confirm presence of "consciousness" in something or rule it out. The only answer to "are LLMs conscious?" is "we don't know". It helps that the whole question is rather meaningles…

Now we have. https://github.com/dmf-archive/IPWT https://dmf-archive.github.io/docs/posts/backpropagation-as-... But you're right, capital only cares about performance. https://dmf-archive.github.io/docs/posts/PoIQ-v2/

This looks to me like the usual "internet schizophrenics inventing brand new theories of everything".

Re: A non-anthropomorphized view of LLMs

#397

Earlier quoted context omitted.

I'm a mind-body dualist and just happened to come across this list, and I think it's an interesting one. #1 we can answer Yes to, #2 through #6 are all strictly unknowable. The best we might be able to claim is some probability distribution that these things may or may not be conscious. The intuitive one looks like 100% chance > P(#2 is conscious) > P(#6) > P(#3) > P(#4) > P(#5) > 0% chance, but the problem is solips…

I think you've fallen into the trap of Descartes' Deus deceptor! Not only is #1 the only question from my list we can definitely answer yes to, but due to this demon this question is actually the only postulate of anything at all that we can answer yes to. All else could be an illusion. Assuming we escape the null space of solipsism, and can reason about anything at all, we can think about what a model might look lik…

To be explicit my P(#) is meant to be the Bayesian probability an observer gives to # being conscious, not the proposition P that # is conscious. It's meant to model Descartes's receptor, as well as disagreement of the kind, "My friend things week 28 fetuses are probably (~% 80%) conscious, and I think they're probably (~20%) not". P(week 28 fetuses) itself is not true or false.

I don't think it's incoherent to make probabilistic claims like this. It might be incoherent to make deeper claims about what laws given the distribution itself. Either way, what I think is interesting is that, if we also think there is such a thing as an amount of consciousness a thing can have, as in the panpsychic view, these two things create an inverse-square law of moral consideration that matches the shape of most people's intuitions oddly well.

For example: Let's say rock is probably not conscious, P(rock) very conscious. A low percentage of a low amount multiplies to a very low expected value, and that matches our intuitions about how much value to give rocks.

Re: A non-anthropomorphized view of LLMs

#398

Earlier quoted context omitted.

Within a single forward pass, but not from one emitted token to another.

What? No. The intermediate hidden states are preserved from one token to another. A token that is 100k tokens into the future will be able to look into the information of the present token's hidden state through the attention mechanism. This is why the KV cache is so big.

KV cache is just that: a cache.

The inference logic of an LLM remains the same. There is no difference in outcomes between recalculating everything and caching. The only difference is in the amount of memory and computation required to do it.

Re: A non-anthropomorphized view of LLMs

#399

Earlier quoted context omitted.

I think you've fallen into the trap of Descartes' Deus deceptor! Not only is #1 the only question from my list we can definitely answer yes to, but due to this demon this question is actually the only postulate of anything at all that we can answer yes to. All else could be an illusion. Assuming we escape the null space of solipsism, and can reason about anything at all, we can think about what a model might look lik…

To be explicit my P(#) is meant to be the Bayesian probability an observer gives to # being conscious, not the proposition P that # is conscious. It's meant to model Descartes's receptor, as well as disagreement of the kind, "My friend things week 28 fetuses are probably (~% 80%) conscious, and I think they're probably (~20%) not". P(week 28 fetuses) itself is not true or false. I don't think it's incoherent to make…

Ah I understand, you're exactly right I misinterpreted the notation of P(#). I was considering each model as assigning binary truth values to the propositions (e.g., physicalism might reject all but Postulate #1, while an anthropocentric model might affirm only #1, #2, and #6), and modeling the probability distribution over those models instead. I think the expected value computation ends up with the same downstream result of distributions over propositions.

By incoherent I was referring to the internal inconsistencies of a model, not the probabilistic claims. Ie a model that denies your own consciousness but accepts the consciousness of others is a difficult one to defend. I agree with your statement here.

Thanks for your comment I enjoyed thinking about this. I learned the estimating distributions approach from the rationalist/betting/LessWrong folks and think it works really well, but I've never thought much about how it applies to something unfalsifiable.

Re: A non-anthropomorphized view of LLMs

#400

Earlier quoted context omitted.

To be explicit my P(#) is meant to be the Bayesian probability an observer gives to # being conscious, not the proposition P that # is conscious. It's meant to model Descartes's receptor, as well as disagreement of the kind, "My friend things week 28 fetuses are probably (~% 80%) conscious, and I think they're probably (~20%) not". P(week 28 fetuses) itself is not true or false. I don't think it's incoherent to make…

Ah I understand, you're exactly right I misinterpreted the notation of P(#). I was considering each model as assigning binary truth values to the propositions (e.g., physicalism might reject all but Postulate #1, while an anthropocentric model might affirm only #1, #2, and #6), and modeling the probability distribution over those models instead. I think the expected value computation ends up with the same downstream…

You're welcome! Probability distributions over inherently unfalsifiable claims is exotic territory at first, but when I see actual philosophers in the wild debate things I often find a back-and-forth of such claims that definitely looks like two people shifting around likelihood values. I take this as evidence that such a process is what's "really" going on when we go one level removed from the arguments and their background assumptions themselves.
Post reply on HN