Earlier quoted context omitted.
That is an impressive improvement, but I think you've missed the most important point: >a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods So DeepMind is to the point where it's a question of whether their generated model or the experimentally determined structure is closest to the actual physical structure.
I don’t think you can say DeepMind could ever be more accurate to the true physical structure since it was built on the same experimental structures that it is being compared to. The limit of accuracy is the experimental data. However, I think we can say that a DeepMind prediction could at least be as good as a new experimental structure.
Let’s say you have 100000 proteins in the training set. Now remove #1 and train on 99999, and then check that it still predicts the same protein result for #1 as the experimental result.
Or remove from training whole sets of proteins by particular teams to find systematic errors made by teams?