Earlier quoted context omitted.
At that point, how does this compare with simply running the model on the CPU?
if you take a model that requires 200GB of VRAM and you run it on the CPU, it requires 200GB of RAM instead. Still unfeasible on consumer hardware. With this approach you can easily do it on 12GB or less of either RAM of VRAM, at several seconds per token instead of tokens per seconds. Very unusable, but certainly interesting!
But I agree, what the OP does is a lot more efficient than this.