I just think about ARM and Intel/AMD. I find it hard to believe that existing GPUs with their memory bandwidth considerations are optimal architecture and that in 20y we will have the same compute model
But the last time I did any work with cuda was a decade ago, so very rough ideas here