Interesting breakdown of a hypothetical attack. The complexity of modern inference engines and the rush to develop them do create some attack surface. But overall it reads more like a "what if" thought experiment. Splitting the GPU and parser is technically doable, but in practice it's trickier, large models run on clusters where the boundaries between components get blurry so defending against this kind of thing wou…
I remember the days when people said opening files like images was safe because it's not an executable. Then people got clever and started exploiting said image libraries with things like numerical overflows.
So, it's really important to ask these questions in a general security sense and think of mitigations before someone finds it's possible and crashes most of the internet.