I have some background but I'm probably not the best person in the world to explain.
The important thing about the transformers model is that it's the first one we have found which keeps unlocking more and more powerful and general cognitive abilities the more resources we throw at it (parameters, exaflops, datasets). I saw some interview with Ilya Sutskever where he says this; it almost certainly won't be the last or best one, but it was the first one.
--
Why was it the first one? How were these guys so clever and other ones couldn't figure it out?
OK so first you need some context. There is a lot of 'Newton standing on the shoulders of giants' going on here. If all of these giants were around in the 1970s, it probably would have been invented then. Heck for all we know something as good was invented in the 1970s but our computers were too smol to benefit from it. This is what John Carmack is currently looking into.
To really notice the scaling benefits of the transformer architecture, they needed to run billion parameter transformer models on linear-algebra-accelerating GPU chips using differentiable programming frameworks. These are some of the giants we are standing on. The research and development pipeline for these amazing GPUs like [thousands of tech companies -> ASML -> TSMC -> NVIDIA] didn't exist until not so long ago. The special properties of transformers wouldn't have been discovered so soon without this hardware stack.
Another giant we are standing on is the differentiable programming linear algebra libraries and frameworks similar to theano or tensorflow or pytorch or jax. They have had things like this under the name 'mathematical programming' like CPLEX but it wasn't as accessible. 'Differentiable programming' is a newish terminology for what used to be called 'automatic differentiation' where 'differentiation' means essentially the same as calculus derivative. Informally it means that these libraries can predict any tiny output effect of any tiny input change as a computationally cheap side-effect of computing the given output, even for complicated calculations. This capability makes optimization easier, in particular it generalizes the 'backpropagation' algorithm of traditional artificial neural networks.
--
What is the transformer model in more nerdy terms.
At one level, it's just a complicatedly parameterized function, where you can fit the parameters by training on data. This viewpoint puts the importance on the computational power applied to training the model with the advantage of differentiable programming. Some will probably guess that the details of the model architecture don't really matter as long as it has sickening amount of parameters and exaflops and dataset. Some version of this viewpoint is probably true in my opinion.
More specifically, the transformer architecture is like a chain of black box differentiable 'soft' lookup tables. The soft queries and keys and values are each lists of floating point numbers (for example a single soft query is a list of numbers, called a vector) and these vectors are stacked into matrices and the soft lookup is processed quickly with fast matrix multiplication tricks. Importantly, all of this is happening inside of a differentiable programming framework which lets you cheaply answer questions about how any small change to the input will affect the output. This capability is used for training, by making trillions of billions of tiny changes to the floating point numbers in the multiplication matrices in the boxes. At the end, the fully trained chain of black box functions can be used to compute a probability distribution over the next token in the message, which lets you generate messages or translate between languages or whatever.