Your skepticism is very healthy, especially in this arena. With video codecs, information theory is ultimately the devil you must answer to at the end of the day. No amount of patents, specifications or algorithmic fantasy can get you away from fundamental constraints.
It seems like the major trade-off being taken right now is along lines of using more memory to buffer additional frames. This can help you in certain scenarios, but in the general case, you cannot ever hope that a prior frame of video has any bearing on future frames of video. It is just exceedingly likely that most frames of video look much like prior frames. So, you can certainly play this game to a point, but you will quickly find yourself on the other end of the bell curve.
You can also play games with ML, but I argue that you are going even further from the fundamental "truth" of your source data with this kind of technique, even if it appears to be a better aesthetic result in isolation of any other concern.
There are also lots of one-off edge cases that have always been impossible to address with any interframe video compression scheme. Just look at the slowmo guys on youtube dump confetti on a 4K camera. No algorithm except for the dumbest intraframe techniques (i.e. JPEG) can faithfully reproduce scenes with information this dense, and usually at the expense of dramatic bandwidth increases.
Bandwidth is cheap and ubiquitous. I say we just use the algorithms that are the fastest and most efficient for our devices. We aren't in 2010 sucking 3G or edge through a straw anymore. Most people can get 20+mbps in their smartphones in decently-populated areas.