>"they struggle to replicate this performance in hospital conditions"
Are there systematic reasons why radiologists in hospitals are inaccurately assessing the AI's output? If the AI models are better than humans in testing novel data then, well, the thing that has changed in a hospital situation compared to the AI-Human testing environment is not the AI, it is the human, under less controlled constraints, additional pressures, workloads, etc. Perhaps the AI's aren't performing as poorly as thought. Perhaps this is why they performed better to begin with. Otherwise, production ML systems are generally not as highly regarded as these models when they perform as significantly below test data sets in production. Some is expected, but "struggle to replicate" implies more.
>"Most tools can only diagnose abnormalities that are common in training data"
Well yes, training on novel examples is one thing. Training on something categorically different is another thing all together. Also there are thresholds of detection. Detecting nothing, or with a a lower confidence, or unknown anomaly, false positive, etc. How much of the inaccuracy isn't wrong, but simply something that is amended or expanded upon when reviewed? Some details here would be useful.
I'm highly skeptical when generalized statements exclude directly relevant information to which an is referring. The few sources provided don't at all cover model accuracy, and the primary factor cited as problematic with AI review, lack of diversity in study composition for women, ethnic variation, children, links to a a meta study that was not at all related to the composition of models and their training data sets.
The article begins as what appears to be a criticism of AI accuracy with the thinness outlined above but then quickly moves on to a "but that's not what radiologists do anyway", and provides a categorical % breakdown of time spent where Personal/Meetings/Meals and some mixture of the others combine to form at least a third that could be categorized as "Time where the human isn't necessary if graphs are being interpreted by models."
I'm not saying there aren't points here, but overall, it simply sounds like the hand-wavy meandering of someone trying to gatekeep a profession whose services could be massively more utilized with more automation, and sure-- perhaps at even higher quality with more radiologists to boot-- but perfect is the enemy of the good etc. on that score, with enormous costs and delays in service in the meantime.