I wouldn't trust a non-radiologist to safely interpret the results of an AI model for radiology, no matter how well that model performs in benchmarks. Similar to how a model that can do "PhD-level research" is of little use to me if I don't have my own PhD in the topic area it's researching for me, because how am I supposed to analyze a 20 page research report and figure out if it's credible or not?
There’s wildly varying levels of quality among these options, even though they could all reasonably be called “PhD-level research.”