It's clear we're searching for the god algorithm of AI, just like physicists are searching for theory of everything. Are transformers the answer though?
I think the transformers architecture, or something very similar with eventually-on-policy time series forecasting in a markov decision process, is the right answer actually and was what I have been trying to make progress on for a long time[1].