I still cannot trust evaluations and benchmarks. How can you prove that the test datasets are truly unseen examples? I think the only way to prove that these models are truly as good as they claim is to wait and see if they are getting adopted in practice.
It would be actually to progress towards the solution of the "black box" problem, the goal of "transparency".
You have to implement a reasoner (etc.), you conceive the best architecture for it - then implement and test it.