Deep learning will be a great way to do compression for sure, both for audio, video and images. I could see that one could download "knowledge sets" for these decompressors. Looking at Google Earth, download the supplemental "knowledge set" for overhead shots of cities and country side. Looking at people, download the supplemental "knowledge set" for faces and clothing, etc. Basically each domain you want to do well…
The "knowledge set" would be a small matrix - the one that allows to decode whatever encoder has put into compressed data. I guess it will be in volume range of 8x8 16-bit floats or so. Maybe three to six such matrices per channel.