It is still not quite at a human level (radiologists prefer human reports in 60% of cases in blinded trials on the best of 3 models, corresponding to an ELO difference of 67 points; the worst model has an ELO difference of 130 compared to humans).
The evaluation was done on 246 X-rays, which is good; it would be better if they were not all chest X-rays (to see generality), and if there were more than only four radiologists from the same country.
The achieved 0.25 clinically-significant errors is impressive, although it is only for the best of 3 models, which can incur bias; averaged across models, it is 0.27, a bit worse than human error. Additionally, I am wondering where they get the human baseline; they state:
> These results are on par with human baselines from prior work [14]
but the citation[1] doesn’t give data in the same format (and its format is honestly better: it indicates that humans make no urgent errors or worse in 64% of reports).
Surprisingly, there is no improvement with model size; the largest model performs the worst.
[1]: https://arxiv.org/pdf/2303.17579