Earlier quoted context omitted.
My favorite argument against SP is zero shot translation. The model learns Japanese-English and Swahili-English and then can translate Japanese-Swahili directly. That shows something more than simple pattern matching happens inside. Besides all arguments based on model capabilities, there is also an argument from usage - LLMs are more like pianos than parrots. People are playing the LLM on the keyboard, making them '…
The translation happens because of token embeddings. We spent a lot of time developing rich embeddings that capture contextual semantics. Once you learn those, translation is “simply” embedding in one language, and disembedding in another. This does not show complex thinking behavior, although there are probably better examples. Translation just isn’t really one of them.
Deep Learning Is Applied Topology
181–190 of 200 posts
Re: Deep Learning Is Applied Topology
#182Re: Deep Learning Is Applied Topology
#183Earlier quoted context omitted.
Anyone who has widely read topics across philosophy, science (physics, biology), economics, politics (policy, power), from practitioners, from original takes, news, etc. ... has managed to understand a tremendous number of relationships due to just words and their syntax. While many of these relationships are related to things we see and do in trivial ways, the vast majority go far beyond anything that can be seen or…
>Am I really dealing in semantics? Or have I just learned the graph-like latent representation for (statistical or reliable) invariant relationships in a bunch of syntax? This and the rest of the comment are philosophical skepticism, and Kant blew this apart back when Hume's "bundle of experience" model of human subjects was considered an open problem in epistemology.
Re: Deep Learning Is Applied Topology
#184Earlier quoted context omitted.
I guess I'll plug my hobby horse: The whole discourse of "stochastic parrots" and "do models understand" and so on is deeply unhealthy because it should be scientific questions about mechanism, and people don't have a vocabulary for discussing the range of mechanisms which might exist inside a neural network. So instead we have lots of arguments where people project meaning onto very fuzzy ideas and the argument does…
Regardless of the mechanism, the foundational 'conceit' of LLMs is that by dumping enough syntax (and only syntax) into a sufficiently complex system, the semantics can be induced to emerge. Quite a stretch, in my opinion (cf. Plato's Cave).
The truth is, syntax and semantics are strongly intertwined and not cleanly separable. A "proof" is merely a syntactically valid string in some formal system.
Re: Deep Learning Is Applied Topology
#185Earlier quoted context omitted.
Well, we disagree fundamentally. And, I applaud the heavy handed use of condescension. Logical propositions ("2+2=4 regardless of my certainty about it") seem a long way from necessary or sufficient to survival for animals. A fuzzy heatmap of "where is prey going" or "How many prey over there" is much closer to necessary and sufficient. The fact that measurements or senses can update those estimates is a long way fro…
You can enumerate, all you wish, all the fuzzy judgements we need to make. This confirms a capacity for uncertain reasoning. It says nothing about the trivial and innumerable ways concepts compose both in content (imagine that A and-also B) and in logical relation (eg., imagine that not A and B). The point of my "condescension" was to point out that people of your position are arguing from ignorance, with confirmatio…
Re: Deep Learning Is Applied Topology
#186Earlier quoted context omitted.
> It's also vastly energetically cheaper just to have (algorithmic) negation. Even if true, that's an argument that it's cheaper to have something, not that it's cheaper to develop it through natural selection. Training time and energy for LLMs shows how energy intensive training to get to the point of grokking/circuit generalization.
It is a matter of empirical fact that we can reason with logical relationships. Thus taking an LLM and it's training as a model of conginition is empriically false. It should be obviously doubly so, since as a model -- as you point out -- it makes trivial aspects of our cognition impossibly expensive to acqurie.
It is a matter of empirical fact that our ability to correctly reason with logical relationships only has high statistical likelihood, not certainty. This looks less like actual logic and more like a probabilistic model of logic.
Re: Deep Learning Is Applied Topology
#187Since this post is based on my 2014 blog post ( https://colah.github.io/posts/2014-03-NN-Manifolds-Topology/ ), I thought I might comment. I tried really hard to use topology as a way to understand neural networks, for example in these follow ups: - https://colah.github.io/posts/2014-10-Visualizing-MNIST/ - https://colah.github.io/posts/2015-01-Visualizing-Representa... There are places I've found the topological per…
Thanks for the follow up. I've been following your circuits thread for several years now. I find the linear representation hypothesis very compelling, and I have a draft of a review for Toy Models of Superposition sitting in my notes. Circuits I find less compelling, since the analysis there feels very tied to the transformer architecture in specific, but what do I know. Re linear representation hypothesis, surely it…
[0] https://www.anthropic.com/research/superposition-memorizatio...
Re: Deep Learning Is Applied Topology
#188Since this post is based on my 2014 blog post ( https://colah.github.io/posts/2014-03-NN-Manifolds-Topology/ ), I thought I might comment. I tried really hard to use topology as a way to understand neural networks, for example in these follow ups: - https://colah.github.io/posts/2014-10-Visualizing-MNIST/ - https://colah.github.io/posts/2015-01-Visualizing-Representa... There are places I've found the topological per…
Re: Deep Learning Is Applied Topology
#189Earlier quoted context omitted.
> Anyone who has widely read topics across philosophy, science (physics, biology), economics, politics (policy, power), from practitioners, from original takes, news, etc. ... has managed to understand a tremendous number of relationships due to just words and their syntax. You're making a slightly different point from the person you're answering. You're talking about the combination of words (with intelligible conte…
I am using syntax in a general form to mean patterns. We are talking about LLMs and the debate seems to be around whether learning about non-verbal concepts through verbal patterns (i.e. syntax that includes all the rules of word use, including constraints reflecting relations between words meaning, but not communication any of that meaning in more direct ways) constitutes semantic understanding or not. In the end, a…
Appears quite bold. What sense-relays inform us about infinity or other mathematical concepts that don't exist physically? Is math-sense its own sense that pulls from something extra-physical?
Doesn't this also go against Chomsky's work, the poverty of stimulus. That it's the recursive nature of language that provides so much linguistic meaning and ability, not sense data, which would be insufficient?