I have to reread it but the model misspecification interpretation (that highly overparameterized models exhibit DD because they reduce misspecification error) is the only thing I've read in this area that makes some sense to me theoretically.
I think DD is a huge issue for a number of fields and is really underappreciated a lot. Without meaning to sound disrespectful, much of this literature seems a little superficial or dismissive, not aware of the broad implications of the claims often being made.
This is because of the ties between information-theory and statistics/modeling. In some sense, at least in the way I've thought about it, the DD seems to imply some kind of violations of fundamental information theory and comes across to me a bit as if someone in chemistry started claiming that some basic laws of thermodynamics in physics didn't apply anymore. Basically, the DD seems to imply that someone can extract more information from a string than the string contains. If you put it this way, it makes no sense, which is why I think this is such a hugely important issue.
On the other hand, the empirical results are there, so figuring out what's going on is worthwhile and I have an open mind.
This paper seems nice with the misspecification angle, because it is realistic and seems to open a path to some interpretations that might not violate some fundamental identities in IT. Misspecification (mismatched coding in IT) can lead to some weird phenomena that's not always intuitive.
Another thing in the paper that's made clear is that DD might not always happen, and it seems informative to figure out when that's the case.
In the background I have to say I'm still skeptical of the empirical breadth of DD. These weird cases of ML failures due to subtle challenge inputs (the example of errors in identifying Obama based on positioning and ties (?) is one example) to me seems like prime examples of overfitting. I still have a hunch that something about the training and test samples relative to the universe of actual intended samples is at play, or the whole phenomenon of DD is misleading because the overfitting problem is really in terms of model flexibility versus data complexity, and not necessarily in terms of number of parameters per se versus sample size.