Live data from Hacker News

A non-anthropomorphized view of LLMs

addxorrol.blogspot.com

261–270 of 432 posts

Re: A non-anthropomorphized view of LLMs

#261

I have the technical knowledge to know how LLMs work, but I still find it pointless to not anthropomorphize, at least to an extent. The language of "generator that stochastically produces the next word" is just not very useful when you're talking about, e.g., an LLM that is answering complex world modeling questions or generating a creative story. It's at the wrong level of abstraction, just as if you were discussing…

I remember Dawkins talking about the "intentional stance" when discussing genes in The Selfish Gene.

It's flat wrong to describe genes as having any agency. However it's a useful and easily understood shorthand to describe them in that way rather than every time use the full formulation of "organisms who tend to possess these genes tend towards these behaviours."

Sometimes to help our brains reach a higher level of abstraction, once we understand the low level of abstraction we should stop talking and thinking at that level.

Re: A non-anthropomorphized view of LLMs

#262

Earlier quoted context omitted.

> My question: how do we know that this is not similar to how human brains work. It is similar to how human brains operate. LLMs are the (current) culmination of at least 80 years of research on building computational models of the human brain.

> It is similar to how human brains operate. Is it? Do we know how human brains operate? We know the basic architecture of them, so we have a map, but we don't know the details. "The cellular biology of brains is relatively well-understood, but neuroscientists have not yet generated a theory explaining how brains work. Explanations of how neurons collectively operate to produce what brains can do are tentative and in…

The part I was referring to is captured in

"The cellular biology of brains is relatively well-understood"

Fundamentally, brains are not doing something different in kind from ANNs. They're basically layers of neural networks stacked together in certain ways.

What we don't know are things like (1) how exactly are the layers stacked together, (2) how are the sensors (like photo receptors, auditory receptors, etc) hooked up?, (3) how do the different parts of the brain interact?, (4) for that matter what do the different parts of the brain actually do?, (5) how do chemical signals like neurotransmitters convey information or behavior?

In the analogy between brains and artificial neural networks, these sorts of questions might be of huge importance to people building AI systems, but they'd be of only minor importance to users of AI systems. OpenAI and Google can change details about how their various transformer layers and ANN layers are connected. The result may be improved products, but they won't be doing anything different from what AIs are doing now in terms the author of this article is concerned about.

Re: A non-anthropomorphized view of LLMs

#263
post #5

The problem with viewing LLMs as just sequence generators, and malbehaviour as bad sequences, is that it simplifies too much. LLMs have hidden state not necessarily directly reflected in the tokens being produced and it is possible for LLMs to output tokens in opposition to this hidden state to achieve longer term outcomes (or predictions, if you prefer). Is it too anthropomorphic to say that this is a lie? To say th…

Author of the original article here. What hidden state are you referring to? For most LLMs the context is the state, and there is no "hidden" state. Could you explain what you mean? (Apologies if I can't see it directly)

Re: A non-anthropomorphized view of LLMs

#264
post #161

The author seems to want to label any discourse as “anthropomorphizing”. The word “goal” stood out to me: the author wants us to assume that we're anthropomorphizing as soon as we even so much as use the word “goal”. A simple breadth-first search that evaluates all chess boards and legal moves, but stops when it finds a checkmate for white and outputs the full decision tree, has a “goal”. There is no anthropomorphizi…

Author here. I am entirely ok with using "goal" in the context of an RL algorithm. If you read my article carefully, you'll find that I object to the use of "goal" in the context of LLMs.

Re: A non-anthropomorphized view of LLMs

#265

I have the technical knowledge to know how LLMs work, but I still find it pointless to not anthropomorphize, at least to an extent. The language of "generator that stochastically produces the next word" is just not very useful when you're talking about, e.g., an LLM that is answering complex world modeling questions or generating a creative story. It's at the wrong level of abstraction, just as if you were discussing…

On the contrary, anthropomorphism IMO is the main problem with narratives around LLMs - people are genuinely talking about them thinking and reasoning when they are doing nothing of that sort (actively encouraged by the companies selling them) and it is completely distorting discussions on their use and perceptions of their utility.

Well "reasoning" refers to Chain-of-Thought and if you look at the generated prompts it's not hard to see why it's called that.

That said, it's fascinating to me that it works (and empirically, it does work; a reasoning model generating tens of thousands of tokens while working out the problem does produce better results). I wish I knew why. A priori I wouldn't have expected it, since there's no new input. That means it's all "in there" in the weights already. I don't see why it couldn't just one shot it without all the reasoning. And maybe the future will bring us more distilled models that can do that, or they can tease out all that reasoning with more generated training data, to move it from dispersed around the weights -> prompt -> more immediately accessible in the weights. But for now "reasoning" works.

But then, at the back of my mind is the easy answer: maybe you can't optimize it. Maybe the model has to "reason" to "organize its thoughts" and get the best results. After all, if you give me a complicated problem I'll write down hypotheses and outline approaches and double check results for consistency and all that. But now we're getting dangerously close to the "anthropomorphization" that this article is lamenting.

Re: A non-anthropomorphized view of LLMs

#266
> The moment that people ascribe properties such as "consciousness" or "ethics" or "values" or "morals" to these learnt mappings is where I tend to get lost. We are speaking about a big recurrence equation that produces a new word, and that stops producing words if we don't crank the shaft.

If that's the argument, then in my mind the more pertinent question is should you be anthropomorphizing humans, Larry Ellison or not.

Re: A non-anthropomorphized view of LLMs

#267

To claim that LLMs do not experience consciousness requires a model of how consciousness works. The author has not presented a model, and instead relied on emotive language leaning on the absurdity of the claim. I would say that any model one presents of consciousness often comes off as just as absurd as the claim that LLMs experience it. It's a great exercise to sit down and write out your own perspective on how con…

Author here. What's the difference, in your perception, between an LLM and a large-scale meteorological simulation, if there is any?

If you're willing to ascribe the possibility of consciousness to any complex-enough computation of a recurrence equation (and hence to something like ... "earth"), I'm willing to agree that under that definition LLMs might be conscious. :)

Re: A non-anthropomorphized view of LLMs

#268

[flagged]

https://rentry.co/2re4t2kx This is what I got pasting the blog post in a prompt asking deepseep to write a reply in a stereotypical hackernews manner. You are about as useful as a LLM as it can replicate your shallow memetics worthless train of thought.

The LLM is right. That’s the problem. It made good points.

Your super intelligent brain couldn’t come up with a retort so you just used an LLM to reinforce my points, making the genius claim that if an LLM came up with even more points that were as valid as mine then I must be just like an LLM?

Like are you even understanding the LLM generated a superior reply? Your saying I’m no different from ai slop then you proceed to show off a 200 iq level reply from an LLM. Bro… wake up, if you didn’t know it was written by an LLM that reply is so good you wouldn’t even know how to respond. It’s beating you.

Re: A non-anthropomorphized view of LLMs

#269
post #251

It still boggles my mind why an amazing text autocompletion system trained on millions of books and other texts is forced to be squeezed through the shape of a prompt/chat interface, which is obviously not the shape of most of its training data. Using it as chat reduces the quality of the output significantly already.

What's your suggested alternative?

In our internal system we use it "as-is" as an autocomplete system; query/lead into terms directly and see how it continues and what it associates with the lead you gave.

Also visualise the actual associative strength of each token generated to confer how "sure" the model is.

LLMs alone aren't the way to AGI or an individual you can talk to in natural language. They're a very good lossy compression over a dataset that you can query for associations.

Re: A non-anthropomorphized view of LLMs

#270
post #233

Earlier quoted context omitted.

I really like that, I think it has the right amount of distance. They don't write, they model writing. We're very used to "all models are wrong, some are useful", "the map is not the territory", etc.

No one was as bothered when we anthropomorphized crud apps simply for the purpose of conversing about "them". "Ack! The thing is corrupting tables again because it thinks we are still using api v3! Who approved that last MR?!" The fact that people are bothered by the same language now is indicative in itself. If you want to maintain distance, pre prompt models to structure all conversations to lack pronouns as betwee…

> You can have the model call you out for referring to the model as existing.

This tickled me. "There ain't nobody here but us chickens".

I have other thoughts which are not quite crystalized, but I think UX might be having an outsized effect here.

Post reply on HN