Live data from Hacker News

Deep Learning Is Applied Topology

theahura.substack.com

171–180 of 200 posts

Re: Deep Learning Is Applied Topology

#171
Aren't manifolds generally task-dependent?

I've been debating whether the data lies on a manifold, or whether the data attributes that are task-relevant (and of our interest) lie on a manifold?

I suspect it is the latter, but I've seen Platonic Representation Hypothesis that seems to hint it is the former.

Re: Deep Learning Is Applied Topology

#172
post #142

Earlier quoted context omitted.

I guess I'll plug my hobby horse: The whole discourse of "stochastic parrots" and "do models understand" and so on is deeply unhealthy because it should be scientific questions about mechanism, and people don't have a vocabulary for discussing the range of mechanisms which might exist inside a neural network. So instead we have lots of arguments where people project meaning onto very fuzzy ideas and the argument does…

Regardless of the mechanism, the foundational 'conceit' of LLMs is that by dumping enough syntax (and only syntax) into a sufficiently complex system, the semantics can be induced to emerge. Quite a stretch, in my opinion (cf. Plato's Cave).

Anyone who has widely read topics across philosophy, science (physics, biology), economics, politics (policy, power), from practitioners, from original takes, news, etc. ... has managed to understand a tremendous number of relationships due to just words and their syntax.

While many of these relationships are related to things we see and do in trivial ways, the vast majority go far beyond anything that can be seen or felt.

What does economics look like? I don't know, but I know as I puzzle out optimums, or expected outcomes, or whatever, I am moving forms around in my head that I am aware of, can recognize and produce, but couldn't describe with any connection to my senses.

The same when seeking a proof for a conjecture in an idiosyncratic algebra.

Am I really dealing in semantics? Or have I just learned the graph-like latent representation for (statistical or reliable) invariant relationships in a bunch of syntax?

Is there a difference?

Don't we just learn the syntax of the visual world? Learning abstractions such as density, attachment, purpose, dimensions, sizes, that are not what we actually see, which is lots of dot magnitudes of three kinds. And even those abstractions benefit greatly from the words other people use describing those concepts. Because you really don't "see" them.

I would guess that someone who was born without vision, touch, smell or taste, would still develop what we would consider a semantic understanding of the world, just by hearing. Including a non-trivial more-than-syntactic understanding of vision, touch, smell and taste.

Despite making up their own internal "qualia" for them.

Our senses are just neuron firings. The rest is hierarchies of compression and prediction based on their "syntax".

Re: Deep Learning Is Applied Topology

#173

Earlier quoted context omitted.

Regardless of the mechanism, the foundational 'conceit' of LLMs is that by dumping enough syntax (and only syntax) into a sufficiently complex system, the semantics can be induced to emerge. Quite a stretch, in my opinion (cf. Plato's Cave).

Anyone who has widely read topics across philosophy, science (physics, biology), economics, politics (policy, power), from practitioners, from original takes, news, etc. ... has managed to understand a tremendous number of relationships due to just words and their syntax. While many of these relationships are related to things we see and do in trivial ways, the vast majority go far beyond anything that can be seen or…

> Anyone who has widely read topics across philosophy, science (physics, biology), economics, politics (policy, power), from practitioners, from original takes, news, etc. ... has managed to understand a tremendous number of relationships due to just words and their syntax.

You're making a slightly different point from the person you're answering. You're talking about the combination of words (with intelligible content, presumably) and the syntax that enables us to build larger ideas from them. The person you're answering is saying that LLM work on the principle that it's possible for intelligence to emerge (in appearance if not in fact) just by digesting a syntax and reproducing it. I agree with the person you're answering. Please excuse the length of the below, as this is something I've been thinking about a lot lately, so I'm going to do a short brain dump to get it off my chest:

The Chinese Room thought experiment --treated by the Stanford Encyclopedia of Philosophy as possibly the single most discussed and debated thought experiment of the latter half of the 20th century -- argued precisely that no understanding can emerge from syntax, and thus by extension that 'strong AI', that really, actually understands (whatever we mean by that) is impossible. So plenty of people have been debating this.

I'm not a specialist in continental philosophy or social thought, but, similarly, it's my understanding that structuralism argued essentially the one can (or must) make sense of language and culture precisely by mapping their syntax. There aren't structulists anymore, though. Their project failed, because their methods don't work.

And, again, I'm no specialist, so take this with a grain of salt, but poststructuralism was, I think, built partly on the recognition that such syntax is artificial and artifice. The content, the meaning, lives somewhere else.

The 'postmodernism' that supplanted it, in turn, tells us that the structuralists were basically Platonists or Manicheans -- treating ideas as having some ideal (in a philosophical sense) form separate from their rough, ugly, dirty, chaotic embodiments in the real world. Postmodernism, broadly speaking, says that that's nonsense (quite literally) because context is king (and it very much is).

So as far as I'm aware, plenty of well informed people whose very job is to understand these issues still debate whether syntax per se confers any understanding whatsoever, and the course philosophy followed in the 20th century seems to militate, strongly, against it.

Re: Deep Learning Is Applied Topology

#174
post #142

Earlier quoted context omitted.

Related to ways of understanding neural networks, I've seen these views expressed a lot, which to me seem like misconceptions: - LLMs are basically just slightly better `n-gram` models - The idea of "just" predicting the next token, as if next-token-prediction implies a model must be dumb (I wonder if this [1] popular response to Karpathy's RNN [2] post is partly to blame for people equating language neural nets with…

I guess I'll plug my hobby horse: The whole discourse of "stochastic parrots" and "do models understand" and so on is deeply unhealthy because it should be scientific questions about mechanism, and people don't have a vocabulary for discussing the range of mechanisms which might exist inside a neural network. So instead we have lots of arguments where people project meaning onto very fuzzy ideas and the argument does…

My favorite argument against SP is zero shot translation. The model learns Japanese-English and Swahili-English and then can translate Japanese-Swahili directly. That shows something more than simple pattern matching happens inside.

Besides all arguments based on model capabilities, there is also an argument from usage - LLMs are more like pianos than parrots. People are playing the LLM on the keyboard, making them 'sing'. Pianos don't make music, but musicians with pianos do. Bender and Gebru talk about LLMs as if they work alone, with no human direction. Pianos are also dumb on their own.

Re: Deep Learning Is Applied Topology

#175
post #142

Earlier quoted context omitted.

I guess I'll plug my hobby horse: The whole discourse of "stochastic parrots" and "do models understand" and so on is deeply unhealthy because it should be scientific questions about mechanism, and people don't have a vocabulary for discussing the range of mechanisms which might exist inside a neural network. So instead we have lots of arguments where people project meaning onto very fuzzy ideas and the argument does…

Regardless of the mechanism, the foundational 'conceit' of LLMs is that by dumping enough syntax (and only syntax) into a sufficiently complex system, the semantics can be induced to emerge. Quite a stretch, in my opinion (cf. Plato's Cave).

> Regardless of the mechanism, the foundational 'conceit' of LLMs is that by dumping enough syntax (and only syntax) into a sufficiently complex system, the semantics can be induced to emerge.

Syntax has dual aspect. It is both content and behavior (code and execution, or data and rules, form and dynamics). This means syntax as behavior can process syntax as data. And this is exactly how neural net training works. Syntax as execution (the model weights and algorithm) processes syntax as data (activations and gradients). In the forward pass the model processes data, producing outputs. In the backward pass it is the weights of the model that become the data to be processed.

When such a self-generative syntactic system is in contact with an environment, in our case the training set, it can encode semantics. Inside the model data is relationally encoded in the latent space. Any new input stands in relation to all past inputs. So data creates its own semantic space with no direct access to the thing in itself. The meaning of a data point is how it stands in relation to all other data points.

Another important aspect is that this process is recursive. A recursive process can't be fully understood from outside. Godel, Turing, Chaitin prove that recursion produces blindspots, that you need to walk the recursive path to know it, you have to be it to know it. Training and inferencing models is such a process

The water carves its banks

The banks channel the water

Which is the true river?

Here, banks = model weights and water = language

Re: Deep Learning Is Applied Topology

#176

Earlier quoted context omitted.

Regardless of the mechanism, the foundational 'conceit' of LLMs is that by dumping enough syntax (and only syntax) into a sufficiently complex system, the semantics can be induced to emerge. Quite a stretch, in my opinion (cf. Plato's Cave).

Anyone who has widely read topics across philosophy, science (physics, biology), economics, politics (policy, power), from practitioners, from original takes, news, etc. ... has managed to understand a tremendous number of relationships due to just words and their syntax. While many of these relationships are related to things we see and do in trivial ways, the vast majority go far beyond anything that can be seen or…

>Am I really dealing in semantics? Or have I just learned the graph-like latent representation for (statistical or reliable) invariant relationships in a bunch of syntax?

This and the rest of the comment are philosophical skepticism, and Kant blew this apart back when Hume's "bundle of experience" model of human subjects was considered an open problem in epistemology.

Re: Deep Learning Is Applied Topology

#177
post #142

Earlier quoted context omitted.

I guess I'll plug my hobby horse: The whole discourse of "stochastic parrots" and "do models understand" and so on is deeply unhealthy because it should be scientific questions about mechanism, and people don't have a vocabulary for discussing the range of mechanisms which might exist inside a neural network. So instead we have lots of arguments where people project meaning onto very fuzzy ideas and the argument does…

My favorite argument against SP is zero shot translation. The model learns Japanese-English and Swahili-English and then can translate Japanese-Swahili directly. That shows something more than simple pattern matching happens inside. Besides all arguments based on model capabilities, there is also an argument from usage - LLMs are more like pianos than parrots. People are playing the LLM on the keyboard, making them '…

> The model learns Japanese-English and Swahili-English and then can translate Japanese-Swahili directly. That shows something more than simple pattern matching happens inside.

The "water story" is a pivotal moment in Helen Keller's life, marking the start of her communication journey. It was during this time that she learned the word "water" by having her hand placed under a running pump while her teacher, Anne Sullivan, finger-spelled the word "w-a-t-e-r" into her other hand. This experience helped Keller realize that words had meaning and could represent objects and concepts.

As the above human experience shows, aligning tokens from different modalities is the first step in doing anything useful.

Re: Deep Learning Is Applied Topology

#178

Earlier quoted context omitted.

Anyone who has widely read topics across philosophy, science (physics, biology), economics, politics (policy, power), from practitioners, from original takes, news, etc. ... has managed to understand a tremendous number of relationships due to just words and their syntax. While many of these relationships are related to things we see and do in trivial ways, the vast majority go far beyond anything that can be seen or…

> Anyone who has widely read topics across philosophy, science (physics, biology), economics, politics (policy, power), from practitioners, from original takes, news, etc. ... has managed to understand a tremendous number of relationships due to just words and their syntax. You're making a slightly different point from the person you're answering. You're talking about the combination of words (with intelligible conte…

I am using syntax in a general form to mean patterns.

We are talking about LLMs and the debate seems to be around whether learning about non-verbal concepts through verbal patterns (i.e. syntax that includes all the rules of word use, including constraints reflecting relations between words meaning, but not communication any of that meaning in more direct ways) constitutes semantic understanding or not.

In the end, all the meaning we have is constructed from the patterns our senses relay to us. We construct meaning from those patterns.

I.e. LLMs may or may not “understand” as well or deeply as we do. But what they are doing is in the same direction.

Re: Deep Learning Is Applied Topology

#179
post #82

Earlier quoted context omitted.

The reason deep learning is alchemy is that none of these deep theories have predictive ability. Essentially all practical models are discovered by trial and error and then "explained" after the fact. In many papers you read a few paragraphs of derivation followed by a simpler formulation that "works better in practice". E.g., diffusion models: here's how to invert the forward diffusion process, but actually we don't…

It's true that there is no directly predictive model of deep learning, and it's also true that there is some trial and error, but it is wrong to say that therefore there is no operating theory at all. I recommend reading Ilyas 30 papers (here's my review of that set: https://open.substack.com/pub/theahura/p/ilyas-30-papers-to-... ) to see how shared intuitions and common threads are clearly developed over the last de…

That is a great list, do you know of something similar that is more recent?

Re: Deep Learning Is Applied Topology

#180
post #142

Earlier quoted context omitted.

I guess I'll plug my hobby horse: The whole discourse of "stochastic parrots" and "do models understand" and so on is deeply unhealthy because it should be scientific questions about mechanism, and people don't have a vocabulary for discussing the range of mechanisms which might exist inside a neural network. So instead we have lots of arguments where people project meaning onto very fuzzy ideas and the argument does…

My favorite argument against SP is zero shot translation. The model learns Japanese-English and Swahili-English and then can translate Japanese-Swahili directly. That shows something more than simple pattern matching happens inside. Besides all arguments based on model capabilities, there is also an argument from usage - LLMs are more like pianos than parrots. People are playing the LLM on the keyboard, making them '…

The translation happens because of token embeddings. We spent a lot of time developing rich embeddings that capture contextual semantics. Once you learn those, translation is “simply” embedding in one language, and disembedding in another.

This does not show complex thinking behavior, although there are probably better examples. Translation just isn’t really one of them.

Post reply on HN