Earlier quoted context omitted.
You could make it much harder to detect by synthesizing a unique video with a DNN and hiding the data using traditional stenography techniques.
I think that video compression might make this not a viable technique. Artifacts would destroy the hidden data, right?
something like this but far more mundane