Earlier quoted context omitted.
>research from Anthropic [1] suggests that structures corresponding to meaning exist inside those bundles of numbers and that there are signs of activity within those bundles of numbers that seem analogous to thought. Can you give some concrete examples? The link you provided is kind of opaque >Amanda Askell [2]. When she talks, she anthropomorphizes LLMs constantly. She is a philosopher by trade and she describes he…
Here are three of the Anthropic research reports I had in mind: https://www.anthropic.com/news/golden-gate-claude Excerpt: “We found that there’s a specific combination of neurons in Claude’s neural network that activates when it encounters a mention (or a picture) of this most famous San Francisco landmark.” https://www.anthropic.com/research/tracing-thoughts-language... Excerpt: “Recent research on smaller models h…
Bag of words, have mercy on us
311–320 of 362 posts
Re: Bag of words, have mercy on us
#312Everyone is out here acting like "predicting the next thing" is somehow fundamentally irrelevant to "human thinking" and it is simply not the case. What does it mean to say that we humans act with intent? It means that we have some expectation or prediction about how our actions will effect the next thing, and choose our actions based on how much we like that effect. The ability to predict is fundamental to our abili…
In the case of LLMs you run into similarities, but they're much more monolithic networks, so the aggregate activations are going to scan across billions of neurons each pass. The sub-networks you can select each pass by looking at a threshold of activations resemble the diverse set of semantic clusters in bio brains - there's a convergent mechanism in how LLMs structure their model of the world and how brains model the world.
This shouldn't be surprising - transformer networks are designed to learn the complex representations of the underlying causes that bring about things like human generated text, audio, and video.
If you modeled a star with a large transformer model, you would end up with semantic structures and representations that correlate to complex dynamic systems within the star. If you model slug cellular growth, you'll get structure and semantics corresponding to slug DNA. Transformers aren't the end-all solution - the paradigm is missing a level of abstraction that fully generalizes across all domains, but it's a really good way to elicit complex functions from sophisticated systems, and by contrasting the way in which those models fail against the way natural systems operate, we'll find better, more general methods and architectures, until we cross the threshold of fully general algorithms.
Biological brains are a computational substrate - we exist as brains in bone vats, connected to a wonderfully complex and sophisticated sensor suite and mobility platform that feeds electrically activated sensory streams into our brains, which get processed into a synthetic construct we experience as reality.
Part of the underlying basic functioning of our brains is each individual column performing the task of predicting which of any of the columns it's connected to will fire next. The better a column is at predicting, the better the brain gets at understanding the world, and biological brains are recursively granular across arbitrary degrees of abstraction.
LLMs aren't inherently incapable of fully emulating human cognition, but the differences they exhibit are expensive. It's going to be far more efficient to modify the architecture, and this may diverge enough that whatever the solution ends up being, it won't reasonably be called an LLM. Or it might not, and there's some clever tweak to things that will push LLMs over the threshold.
Re: Bag of words, have mercy on us
#313Earlier quoted context omitted.
LLMs are compression and prediction. The most efficient way to (lossfully) compress most things is by actually understanding them. Not saying LLMs are doing a good job of that, but that is the fundamental mechanism here.
Where’s the proof that efficient compression results in “understanding”? Is there a rigorous model or theorem, or did you just make this up?
If you're interested in why compression is like understanding in many ways, I'd suggest reading through the wikipedia article on Kolmogorov complexity.
Re: Bag of words, have mercy on us
#314Earlier quoted context omitted.
> How did these 2 models do it if not actually using language like a thinking agent? By having a gazillion of other, almost identical pictures of kids in parks in their training data.
Not pictures with this composition, same jacket, etc - yes, there are images but they are different, this fits like a key in the lock to the original
Re: Bag of words, have mercy on us
#315Re: Bag of words, have mercy on us
#316Everyone is out here acting like "predicting the next thing" is somehow fundamentally irrelevant to "human thinking" and it is simply not the case. What does it mean to say that we humans act with intent? It means that we have some expectation or prediction about how our actions will effect the next thing, and choose our actions based on how much we like that effect. The ability to predict is fundamental to our abili…
They prove to have some useful utility to me regardless.
Re: Bag of words, have mercy on us
#317Earlier quoted context omitted.
Like a car engine or a combine. The problem isn't the effectiveness of the tool for its purpose, it's the religion around it.
We seem to be talking past one another. All I was talking about was the facts of how these systems perform, without any reverence about it at all. But to your point, I do see a lot of people very emotionally and psychologically committed to pointing out how deeply magical humans are, and how impossible we are to replicate in silicon. We have a religion about ourselves; we truly do have main character syndrome. It's w…
This a straw man, the question isn't if this is possible or not (this is an open question), it's about whether or not we are already here, and the answer is pretty straightforward: no we aren't. (And the current technology isn't going to bring us anywhere near that)
Re: Bag of words, have mercy on us
#318Re: Bag of words, have mercy on us
#319Earlier quoted context omitted.
I'm definitely a stream of words. My "abstract thoughts" are a stream of words too, they just don't get sounded out. Tbf I'd rather they weren't there in the first place. But bodies which refuse to harbor an "interiority" are fast-tracked to destruction because they can't suf^W^W^W be productive. Funny movie scene from somewhere. The sergeant is drilling the troops: "You, private! What do you live for!", and expects…
Then what are non-human animals doing?
Re: Bag of words, have mercy on us
#320Earlier quoted context omitted.
It's been repeated a huge number of time since, and widely debated. When Libet first did the experiment it was only like 200ms before the mind become consciously aware of the decision. More recent studies have shown they can predict actions up to 7-10 seconds before the subject is aware of having made a decision. It's pretty hard to argue that you're really "free" to make a different decision if your body knew which…
"I conducted an experiment where I instructed experienced drivers to follow a path in a parking lot laid out with traffic cones, and found that we were able to predict the trajectory of the car with greater than 60% accuracy. Therefore drivers do not have free will to just dodge the cones and drive arbitrarily from the start to the finish." Clearly, that conclusion would be patently absurd to draw from that experimen…
Certainly we can come up with some alternative theories (like "free will") to explain it all away, but the simplest (therefore most likely correct) answer is just that we're basically statistical state machines and are as deterministic as a similar computational system.
To be clear, I'm not saying that metacognition doesn't exist. Just that I've never seen any reason to believe it's very different from current thinking models that just feed an output back in as another input.
[0] - https://home.csulb.edu/~cwallis/382/readings/482/nisbett%20s...