Earlier quoted context omitted.
I was stumped by that as well. It spends half the text explaining the well-known parts (surface models) and then mentions the actual contribution completely in passing: > So first his system renders a 3D reconstruction of the 2D image in very low resolution [...] Next, do a higher-resolution render of the area you kept. Where apparently the author seems to think producing a 3D reconstruction of a 2D object is trivial…
That's not what the paper is about. The abstract quoted by the GP shows that it has been done before. Probably only very recently so it's going to be a surprise to a lot of readers, but perhaps old news measured in machine learning years.
Oh my god, I sound like my father.