Live data from Hacker News

International evaluation of an AI system for breast cancer screening

nature.com

1–10 of 28 posts

Re: International evaluation of an AI system for breast cancer screening

#2
> Screening mammography aims to identify breast cancer at earlier stages of the disease, when treatment can be more successful. Despite the existence of screening programmes worldwide, the interpretation of mammograms is affected by high rates of false positives and false negatives. Here we present an artificial intelligence (AI) system that is capable of surpassing human experts in breast cancer prediction. [...]

> In an independent study of six radiologists, the AI system outperformed all of the human readers: the area under the receiver operating characteristic curve (AUC-ROC) for the AI system was greater than the AUC-ROC for the average radiologist by an absolute margin of 11.5%. We ran a simulation in which the AI system participated in the double-reading process that is used in the UK, and found that the AI system maintained non-inferior performance and reduced the workload of the second reader by 88%. This robust assessment of the AI system paves the way for clinical trials to improve the accuracy and efficiency of breast cancer screening.

So, there you have it: AI not "either/or" humans, but both, in conjunction, as a composition of the best of both worlds.

At the very least, that's how civilization will massively and intimately introduce true assistant AI.

It's also somewhat counter-intuitive to think that the most specialized tasks are the low hanging fruits; i.e. that the "difficult" to us, culminating years of training and experience for humans (e.g. how to read a medical scan) may be, per its natural advantages (like speed and parallelism), "easy" to the machine.

That space (where machine expertise is cheaper than human) roughly maps to the immense value attributed to the rise of industrial-age narrow AI; therein lies not a way to replace humans — we never did that in history, merely destroyed jobs to create ever more — but rather to augment ourselves once more to whole new levels of performance.

Anything more than this is AGI-level, science-fiction so far — and there's not even a shred of evidence that it's theoretically a sure thing, possible in the first place. Which is not to say that AI safety research isn't extremely important even for the narrow kind (manipulation comes to mind), but we shouldn't go as far as to bet future economic growth on its existence. Like fusion or interstellar travel, we just don't know. Yet, and for the foreseeable future, because scale.

Re: International evaluation of an AI system for breast cancer screening

#3

> Screening mammography aims to identify breast cancer at earlier stages of the disease, when treatment can be more successful. Despite the existence of screening programmes worldwide, the interpretation of mammograms is affected by high rates of false positives and false negatives. Here we present an artificial intelligence (AI) system that is capable of surpassing human experts in breast cancer prediction. [...] >…

Exactly this. This is where I see AI possibly going: To be a complimentary tool or second pair of eyes to speed up the work for the professionals rather than replacing them. I also see this research as a very positive step forward for using AI for good and especially bringing highly accurate results that can used as a aid for health professionals.

However, given that this research used a deep learning (DL) based AI system in the medical industry, there are still questions around this AI system explaining itself and its internal decision process for the sake of transparency, which will almost be ignored in other news reporting sites and will focus only on the accuracy. DL-based AI systems will still be a concern towards both patients and clinicians and I would expect this to be a focus point in the future, despite the welcoming results which is still very interesting anyways.

Other than the transparency issues behind the AI system, I'd say this is a great start into the new decade for AI.

Re: International evaluation of an AI system for breast cancer screening

#4
I’m not a huge fan of turning the BI-RADS classification scheme into a ROC curve. From what I’ve seen, BI-RADS is something like a yes / no / maybe scheme for mammograms. I don’t think it was designed to be treated like a test score, so using it to generate a ROC curve feels like an unfair comparison between the AI system and current clinical practice.

What they’re doing is interesting, but it’s still very academic. I have little doubt that eventually some sort of AI system will benefit clinical practice, but based on the sheer number of studies that fail to make it over the line, I’m not sure I have high hopes for this one. Why they’ve done so far is the equivalent of “It works in vitro...”

Re: International evaluation of an AI system for breast cancer screening

#5
> Notably, the additional cancers identified by the AI system tended to be invasive rather than in situ disease.

This is probably due to invasive cancers being more common (~80%) than in situ cancers. I am not sure why this natural explanation was not suggested.

Re: International evaluation of an AI system for breast cancer screening

#6

> Screening mammography aims to identify breast cancer at earlier stages of the disease, when treatment can be more successful. Despite the existence of screening programmes worldwide, the interpretation of mammograms is affected by high rates of false positives and false negatives. Here we present an artificial intelligence (AI) system that is capable of surpassing human experts in breast cancer prediction. [...] >…

How many years did centaurs reign supreme over pure AI in chess? 5-10 maybe? This "both" stuff is just a temporary stop on the way to meat obsolescence.

Re: International evaluation of an AI system for breast cancer screening

#7

I’m not a huge fan of turning the BI-RADS classification scheme into a ROC curve. From what I’ve seen, BI-RADS is something like a yes / no / maybe scheme for mammograms. I don’t think it was designed to be treated like a test score, so using it to generate a ROC curve feels like an unfair comparison between the AI system and current clinical practice. What they’re doing is interesting, but it’s still very academic.…

As I understand, for comparison, they also turned AI system output to BI-RADS class and then back to ROC curve.

Re: International evaluation of an AI system for breast cancer screening

#9
post #7

I’m not a huge fan of turning the BI-RADS classification scheme into a ROC curve. From what I’ve seen, BI-RADS is something like a yes / no / maybe scheme for mammograms. I don’t think it was designed to be treated like a test score, so using it to generate a ROC curve feels like an unfair comparison between the AI system and current clinical practice. What they’re doing is interesting, but it’s still very academic.…

As I understand, for comparison, they also turned AI system output to BI-RADS class and then back to ROC curve.

Huh, where does it say that in the article? I don’t think I spotted that.

All the same, it feels sort of beside the point to me. It just doesn’t feel right to take a medical diagnostic tool - whose intended purpose is for communication among doctors - and treat it as a test score. That’s just... not what it was designed for.

Re: International evaluation of an AI system for breast cancer screening

#10
post #7

Earlier quoted context omitted.

As I understand, for comparison, they also turned AI system output to BI-RADS class and then back to ROC curve.

Huh, where does it say that in the article? I don’t think I spotted that. All the same, it feels sort of beside the point to me. It just doesn’t feel right to take a medical diagnostic tool - whose intended purpose is for communication among doctors - and treat it as a test score. That’s just... not what it was designed for.

In page 3, "Readers rated each case using the forced BI-RADS scale, and BI-RADS scores were compared to ground-truth outcomes to fit an ROC curve for each reader. The scores of the AI system were treated in the same manner (Fig. 3)."

This isn't as clear as I want it to be, but Fig. 3 shows both "AI system" and "AI system (non-parametric)" ROC curve. My understanding is that the former is fit from discrete BI-RADS class, and the latter is "raw" output.

Post reply on HN