Live data from Hacker News

What are transformer models and how do they work?

txt.cohere.ai

61–70 of 114 posts

Re: What are transformer models and how do they work?

#61
post #59
post #49

Oh, come on - this article is from three days ago, but it starts with "Transformers are a new development in machine learning". Transformers have been around for six years now, that's an eternity when you consider how fast this field is moving.

Deep learning (multilayer perceptrons) was 1965... The timeline is pretty long.

Maybe on a theoretical level, but of course you will agree that what we mean by deep learning today has only become possible with the availability of sufficient computational power (and Hinton's work around the mid-2000's).

Re: What are transformer models and how do they work?

#62

Skimming it, there are a few things about this explanation that rub me just slightly the wrong way. 1. Calling the input token sequence a "command". It probably only makes sense to think of this as a "command" on a model that's been fine-tuned to treat it as such. 2. Skipping over BPE as part of tokenization - but almost every transformer explainer does this, I guess. 3. Describing transformers as using a "word embed…

The positional embedding can be thought of: in the same way you can hear two pieces of music overlaid on each other, you can add both the vocab and pos embedding and it’s able to pick them apart. If you asked yourself to identify when someone’s playing a high note or low note (pos embedding) and whether they’re playing Beethoven or Lady Gaga (vocab embedding) you could do it. That’s why it’s additive and why it would…

The visualisation here may be helpful.

https://github.com/tensorflow/tensor2tensor/issues/1591

Re: What are transformer models and how do they work?

#63

Skimming it, there are a few things about this explanation that rub me just slightly the wrong way. 1. Calling the input token sequence a "command". It probably only makes sense to think of this as a "command" on a model that's been fine-tuned to treat it as such. 2. Skipping over BPE as part of tokenization - but almost every transformer explainer does this, I guess. 3. Describing transformers as using a "word embed…

> Skipping over BPE as part of tokenization

Well, there are other methods in use. See ByT5, for example.

Re: What are transformer models and how do they work?

#64
post #24

The more I learn about the technical details of how ML systems are implemented, the more I feel that those details obscure , rather than illuminate, what is actually going on. It's as if we were trying to understand human ethics by looking at neurotransmitters or synapses in the brain. These structures seem way too low-level to actually explain the interesting stuff. What I hear is "something something transformer au…

> Where is the connection between computational details and the model's high-level behavior? Do we even know?

This is an active area of study ("mechanistic interpretability") and it's very early days. For instance here's a paper I read recently that tries to explain how a very simple transformer learns how to do modular arithmetic: https://arxiv.org/abs/2301.05217

Curious what interesting results people are aware of in this area.

Re: What are transformer models and how do they work?

#65
post #44

Earlier quoted context omitted.

This is called emergent behavior, and we don't know how it happens with organic brains, minds, and neurons either. It's actually pretty amazing that it's happening at all with computers, since neural nets are such simple, high level abstractions compared to how the brain works. It's possible that all the tremendous complexity of organic systems isn't actually necessary for intelligence or consciousnes, which is simil…

> It's possible that all the tremendous complexity of organic systems isn't actually necessary for intelligence or consciousnes, which is similarly surprising. My guess is most of that complexity is necessary for efficiency , not for basic function. Biological systems are unimaginably efficient at almost everything they do. The information storage density of DNA is within 1-2 orders of magnitude of the upper limit im…

Efficiency and also redundancy. Brains and bodies have incredible tolerance to damage.

Re: What are transformer models and how do they work?

#66
post #8

Thanks. ML noob here. I liked the insight that attention adds context, by modifying the distance in an embedding.

The way the article presents this is misleading. The attention mechanism builds a new vector as a linear combination of other vectors, but after the first layer these have also all been altered by passing through a transformer layer so it makes less sense to talk about "other tokens" in most cases (it becomes increasingly inaccurate the deeper into the model you go). It's also not really moving closer so much as addi…

You still have one vector per token, that's what they meant, also the fact that the vector associated with each token will ultimately be used to predict the next token, once again showing that it makes sense to talk about other tokens even though they're being transformed inside the model.

Re: What are transformer models and how do they work?

#67
post #61
post #59

Earlier quoted context omitted.

Deep learning (multilayer perceptrons) was 1965... The timeline is pretty long.

Maybe on a theoretical level, but of course you will agree that what we mean by deep learning today has only become possible with the availability of sufficient computational power (and Hinton's work around the mid-2000's).

There are lots of important moments I guess, I would credit the main break in practicality to the ~1990 Swiss/German stuff: https://people.idsia.ch/~juergen/deep-learning-miraculous-ye...

But still even the early perceptrons stuff was applied research, I wouldn't call it purely theoretical by any means.

Re: What are transformer models and how do they work?

#68
post #24

The more I learn about the technical details of how ML systems are implemented, the more I feel that those details obscure , rather than illuminate, what is actually going on. It's as if we were trying to understand human ethics by looking at neurotransmitters or synapses in the brain. These structures seem way too low-level to actually explain the interesting stuff. What I hear is "something something transformer au…

This is called emergent behavior, and we don't know how it happens with organic brains, minds, and neurons either. It's actually pretty amazing that it's happening at all with computers, since neural nets are such simple, high level abstractions compared to how the brain works. It's possible that all the tremendous complexity of organic systems isn't actually necessary for intelligence or consciousnes, which is simil…

"It's possible that all the tremendous complexity of organic systems isn't actually necessary for intelligence or consciousnes, which is similarly surprising."

At the hardware level it's not at all surprising; consider cells, dna, proteins, and so on making up muscles. Compared to a magnet and some coils of copper.

But I think you mean the 'architectural' or connectome complexity of the brain compared to GPT, and I agree it's surprising that such a simple model as GPT is so capable.

Re: What are transformer models and how do they work?

#69
post #67
post #61

Earlier quoted context omitted.

Maybe on a theoretical level, but of course you will agree that what we mean by deep learning today has only become possible with the availability of sufficient computational power (and Hinton's work around the mid-2000's).

There are lots of important moments I guess, I would credit the main break in practicality to the ~1990 Swiss/German stuff: https://people.idsia.ch/~juergen/deep-learning-miraculous-ye... But still even the early perceptrons stuff was applied research, I wouldn't call it purely theoretical by any means.

I agree about your notion of there being lots of important moments in history, and Schmidhuber's contributions are not small by any means. Yet, 1990 was not the year when Deep Learning took off.

Re: What are transformer models and how do they work?

#70

> In short, text embeddings send every piece of text to a vector (a list) of numbers. What does "sent to" mean? Is that baby-talk for "mapped to"?

I sometimes wish why don't mathematical equations come with simple visual examples to help students build mental models (an example here [0]). It is difficult to parse meaning behind equations if written in terse language.

[0] Tai-Danae Bradley: "Entropy as an Operad Derivation" https://www.youtube.com/watch?v=_cAEfQQcELA

Post reply on HN