Live data from Hacker News

Deep Learning Is Applied Topology

theahura.substack.com

181–190 of 200 posts

Re: Deep Learning Is Applied Topology

#181

Earlier quoted context omitted.

My favorite argument against SP is zero shot translation. The model learns Japanese-English and Swahili-English and then can translate Japanese-Swahili directly. That shows something more than simple pattern matching happens inside. Besides all arguments based on model capabilities, there is also an argument from usage - LLMs are more like pianos than parrots. People are playing the LLM on the keyboard, making them '…

The translation happens because of token embeddings. We spent a lot of time developing rich embeddings that capture contextual semantics. Once you learn those, translation is “simply” embedding in one language, and disembedding in another. This does not show complex thinking behavior, although there are probably better examples. Translation just isn’t really one of them.

This is also the problem I have with John Searle’s Chinese room

Re: Deep Learning Is Applied Topology

#183
post #176

Earlier quoted context omitted.

Anyone who has widely read topics across philosophy, science (physics, biology), economics, politics (policy, power), from practitioners, from original takes, news, etc. ... has managed to understand a tremendous number of relationships due to just words and their syntax. While many of these relationships are related to things we see and do in trivial ways, the vast majority go far beyond anything that can be seen or…

>Am I really dealing in semantics? Or have I just learned the graph-like latent representation for (statistical or reliable) invariant relationships in a bunch of syntax? This and the rest of the comment are philosophical skepticism, and Kant blew this apart back when Hume's "bundle of experience" model of human subjects was considered an open problem in epistemology.

Can you get into more detail and share some links? Inquiring minds want to know

Re: Deep Learning Is Applied Topology

#184
post #142

Earlier quoted context omitted.

I guess I'll plug my hobby horse: The whole discourse of "stochastic parrots" and "do models understand" and so on is deeply unhealthy because it should be scientific questions about mechanism, and people don't have a vocabulary for discussing the range of mechanisms which might exist inside a neural network. So instead we have lots of arguments where people project meaning onto very fuzzy ideas and the argument does…

Regardless of the mechanism, the foundational 'conceit' of LLMs is that by dumping enough syntax (and only syntax) into a sufficiently complex system, the semantics can be induced to emerge. Quite a stretch, in my opinion (cf. Plato's Cave).

Curious what you make of symbolic mathematics, then - in particular, systems like Mathematica which can produce true and novel mathematical facts by pure syntactic manipulation.

The truth is, syntax and semantics are strongly intertwined and not cleanly separable. A "proof" is merely a syntactically valid string in some formal system.

Re: Deep Learning Is Applied Topology

#185

Earlier quoted context omitted.

Well, we disagree fundamentally. And, I applaud the heavy handed use of condescension. Logical propositions ("2+2=4 regardless of my certainty about it") seem a long way from necessary or sufficient to survival for animals. A fuzzy heatmap of "where is prey going" or "How many prey over there" is much closer to necessary and sufficient. The fact that measurements or senses can update those estimates is a long way fro…

You can enumerate, all you wish, all the fuzzy judgements we need to make. This confirms a capacity for uncertain reasoning. It says nothing about the trivial and innumerable ways concepts compose both in content (imagine that A and-also B) and in logical relation (eg., imagine that not A and B). The point of my "condescension" was to point out that people of your position are arguing from ignorance, with confirmatio…

Hey it's possible you're absolutely right, but our arguments have about equal value (none) because they are both presented as personal opinions with zero external support. I just had the decency to admit that up front.

Re: Deep Learning Is Applied Topology

#186

Earlier quoted context omitted.

> It's also vastly energetically cheaper just to have (algorithmic) negation. Even if true, that's an argument that it's cheaper to have something, not that it's cheaper to develop it through natural selection. Training time and energy for LLMs shows how energy intensive training to get to the point of grokking/circuit generalization.

It is a matter of empirical fact that we can reason with logical relationships. Thus taking an LLM and it's training as a model of conginition is empriically false. It should be obviously doubly so, since as a model -- as you point out -- it makes trivial aspects of our cognition impossibly expensive to acqurie.

> It is a matter of empirical fact that we can reason with logical relationships

It is a matter of empirical fact that our ability to correctly reason with logical relationships only has high statistical likelihood, not certainty. This looks less like actual logic and more like a probabilistic model of logic.

Re: Deep Learning Is Applied Topology

#187
post #98

Since this post is based on my 2014 blog post ( https://colah.github.io/posts/2014-03-NN-Manifolds-Topology/ ), I thought I might comment. I tried really hard to use topology as a way to understand neural networks, for example in these follow ups: - https://colah.github.io/posts/2014-10-Visualizing-MNIST/ - https://colah.github.io/posts/2015-01-Visualizing-Representa... There are places I've found the topological per…

Thanks for the follow up. I've been following your circuits thread for several years now. I find the linear representation hypothesis very compelling, and I have a draft of a review for Toy Models of Superposition sitting in my notes. Circuits I find less compelling, since the analysis there feels very tied to the transformer architecture in specific, but what do I know. Re linear representation hypothesis, surely it…

I was going to comment the same about the Superposition hypothesis [0], when the OP comment (edit: Update: The OP commenter is (as pointed by other HN comments, the cofounder of Anthropic) behind the Superposition research) mentioned about "I've had a lot more success with: * The linear representation hypothesis - The idea that "concepts" (features) correspond to directions in neural networks", as this concept-per-NN-feature idea seems too "basic" to explain some of the learning which NNs can do on datasets. On one of our custom trained neural network models (not LLM, but audio-based and currently proprietary) we noticed the same of the ML model being able to "overfit" on a large amount of data despite not many few parameters relative to the size of the dataset (and that too with dropout in early layers).

[0] https://www.anthropic.com/research/superposition-memorizatio...

Re: Deep Learning Is Applied Topology

#188
post #98

Since this post is based on my 2014 blog post ( https://colah.github.io/posts/2014-03-NN-Manifolds-Topology/ ), I thought I might comment. I tried really hard to use topology as a way to understand neural networks, for example in these follow ups: - https://colah.github.io/posts/2014-10-Visualizing-MNIST/ - https://colah.github.io/posts/2015-01-Visualizing-Representa... There are places I've found the topological per…

[deleted]

Re: Deep Learning Is Applied Topology

#189

Earlier quoted context omitted.

> Anyone who has widely read topics across philosophy, science (physics, biology), economics, politics (policy, power), from practitioners, from original takes, news, etc. ... has managed to understand a tremendous number of relationships due to just words and their syntax. You're making a slightly different point from the person you're answering. You're talking about the combination of words (with intelligible conte…

I am using syntax in a general form to mean patterns. We are talking about LLMs and the debate seems to be around whether learning about non-verbal concepts through verbal patterns (i.e. syntax that includes all the rules of word use, including constraints reflecting relations between words meaning, but not communication any of that meaning in more direct ways) constitutes semantic understanding or not. In the end, a…

> In the end, all the meaning we have is constructed from the patterns our senses relay to us. We construct meaning from those patterns.

Appears quite bold. What sense-relays inform us about infinity or other mathematical concepts that don't exist physically? Is math-sense its own sense that pulls from something extra-physical?

Doesn't this also go against Chomsky's work, the poverty of stimulus. That it's the recursive nature of language that provides so much linguistic meaning and ability, not sense data, which would be insufficient?

Post reply on HN