Earlier quoted context omitted.
Tracking is not the method they're using for colourizing, but the other way around. Your linked paper has no tracking.
I realize that. They are for some reason saying they can track things with their colorization, when their colorization is extremely unimpressive as well as the tracking that results from using it. There is no reason colorization needs to happen to do the tracking anyway. The tracking is unimpressive and now indirect. This isn't some sort of epiphany they've discovered, they are just reinventing video image segmentati…
As the paper itself states, the tracking results are not the absolute state of the art, but they are in the same ballpark, and more importantly, learned without supervision - just watching video. This makes it easier to train on whatever dataset you might have lying around, and more importantly, it's a clever, simple idea that can be improved on and adapted for different tasks.
(Disclaimer: Authors are acquaintances of mine.)