GNNs trained by backprop are making many of the same mistakes that LSTMs did: solve one important problem(exploding/vanishing gradients), but introduce a bunch of hyperparameters that bring along their own set of problems. GRUs are successful in my opinion, because they remove some of those tunable parameters.
Of course as the post suggested, not being able to tune something, often loses out against some more tuned and curated solution. GNNs being the newest version, having tons of parameters to tweek. In the end do we get a better, ultimately more generalizable solution? Or do we just get more hyperparameters to tune, and spend more time and money for a modest, unrepeatable gain.