Earlier quoted context omitted.
That's the point. With the old models they all failed to produce a wine glass that is completley to the brim full. Because you can't find that a lot in the data they used for training.
Imagine if they just actually trained the model on a bunch of photographs of a full glass of wine, knowing of this litmus test
So maybe training for litmus tests isn’t the worst strategy in the absence of another entire internet of training data…