Earlier quoted context omitted.
It's not the editor's job to replicate the results. Imagine the same requirements in medical research, this would be crazy. Editors accept a paper based on the quality of the paper and the claim. Then the paper needs to be replicated a few times by independent research teams to be considered valid. This is the step we are probably missing in AI. Any published paper (and even unpublished, on arxiv) are considered vali…
> It's not the editor's job to replicate the results. They obviously meant the reviewers. Not everyone is intimately familiar with academic terminology.
Maybe a solution would be a platform, like a CI for machine learning where authors would send their codes, and the CI would run it for them, and make the results public. Then everyone could check the code and results, reviewers included.