Very interesting. This makes me wonder if a similar technique can be used for compression?
In practice, even the hyper-efficient compression algorithms used in something like zpaq tend to use only very small shallow predictive neural networks because no one wants to wait days for their data to be compressed or ship around big neural nets as part of their archives, so it's more of an information-theoretic curiosity. Few enough people will even use 'xz'.