Factoring is something we do to distributions, which we then parameterize with our neural net.
An autoregressive model pertains to a certain factorisation of a joint distribution. And if we choose this factorisation it may have consequences for our training and sampling etc.
But to say "most models on factor generation".. I dont quite get this.
This would be better work if it didnt claim it was a new pretraining axis also. Its another way of scaling compute right?
If it were a new axis, IWAE would have an equal claim to it (as others have pointed out).