This is confusing, using the semantic vectors arithmetic of embeddings is not very relevant to transformers and its completely missing the word 'attention'. I don't think transformers are that difficult to explain to people , but it is hard to explain "why" they work. But i think it's important for everyone to look under the hood and know that there are no demons underneath.
I was trying to keep the article at a level that everyone understands, from middle school up. I thought about going a bit deeper in the structure and mentioning attention, but my problem is that the intuitive concept of "attention" is quite different from the mathematical reality of an attention layer, and I'm sure I would have lost quite a few people there. It's always a trade-off :)
I struggle to understand why this thing works the way it does. It's possible that Vaswani et al. have made one of the greatest discoveries of this century that solved the language problem in an unintuitive, and yet very unappreciated way. It's also possible that there are other architectures that can simulate the same level of intelligence with such large numbers of parameters.
I think you re right that it's not intuitive, it's like basic arithmetic is laughing at us