Live data from Hacker News

AI models that predict disease are not as accurate as reports might suggest

scientificamerican.com

141–150 of 162 posts

Re: AI models that predict disease are not as accurate as reports might suggest

#141
post #36

Earlier quoted context omitted.

Do you agree that it’s ok to pose a question whenever you don’t understand?

I’m not sure where you got this form of communication where you respond to everything with a question, and I assume you mean well, but it comes across as patronizing and de-humanizing to try to follow these “rules to winning arguments passively”, or whatever it is. Indeed, the confusion here is (I think) because your first comment > Sorry for asking, but how is this relevant to the article? Sounds accusatory. Please…

The basic idea of that kind of question is to find the minimal place of agreement. And then understand where one deviates.

Going back the path of arguments to common ground if you will. It works quite well in my experience if you’re interested in genuine discussion.

PS: how something “sounds” is really difficult to say in a written medium. It might say more about the reader than the writer.

Re: AI models that predict disease are not as accurate as reports might suggest

#142

Earlier quoted context omitted.

This has to be intentional no?

The problem is quite subtle, though obvious in retrospect. I've seen a paper from a separate, academic, research group make similar model with the exact same problem. The problem would, however, have been clear, if the model was compared to simply using the current mean blood pressure (MAP) as a predictor of hypotension, because MAP is the problematic predictor variable. Instead, the model was only compared to short-…

Hm, reading the linked tweets the problem seems like a big screaming red target on the side of a white barn, not a feature engineering subtlety. It seems like the typical case of the drunk guy looking for his keys under the streetlight. (Having insufficient data, and comparing the model to an arbitrarily picked one that just happens to be even worse. And then everyone including the FDA patting them on the back.)

Re: AI models that predict disease are not as accurate as reports might suggest

#143
post #38

Earlier quoted context omitted.

are you really sure the doctors are doing a better job when they go through the motions of incorporating a wide range of data? Or do we just convince ourselves they're better? I suspect we massively underestimate the amount of misdiagnosis due to incorrect analysis of data using fairly naive medical mental models of disease.

Not this ignorant comment again. AI will replace software engineers long before it replaces doctors. There is an arrogant ignorance of what doctors do that always shows up in comments when topics like this pop up. And yes I'm a physician and MLE. So i understand both worlds clearly

Well look at that! If it isn’t another member of the medical mafia on HN.

It is kind of funny to see comments complain about the lack of perfect sensitivity and specificity of their physicians.

We complain about the same thing from the the various ML techniques in radiology which currently are pitiful and a gigantic waste of time and money. When I went into rads I was pretty worried about ML - not anymore.

I’m hoping this upcoming recession will dry out a lot of institutional use of ML. In radiology it’s not that helpful and there’s no technical fee increase for it. But you can advertise with it i guess? Commercials with lasers, robots, and AI with pleasant voice overs about cutting edge techniques and getting the care you need in the 21st century and blah blah blah

Re: AI models that predict disease are not as accurate as reports might suggest

#144
post #142

Earlier quoted context omitted.

The problem is quite subtle, though obvious in retrospect. I've seen a paper from a separate, academic, research group make similar model with the exact same problem. The problem would, however, have been clear, if the model was compared to simply using the current mean blood pressure (MAP) as a predictor of hypotension, because MAP is the problematic predictor variable. Instead, the model was only compared to short-…

Hm, reading the linked tweets the problem seems like a big screaming red target on the side of a white barn, not a feature engineering subtlety. It seems like the typical case of the drunk guy looking for his keys under the streetlight. (Having insufficient data, and comparing the model to an arbitrarily picked one that just happens to be even worse. And then everyone including the FDA patting them on the back.)

I'm glad that you seem to get the severity! I'm just hesitant to ascribe malice.

Re: AI models that predict disease are not as accurate as reports might suggest

#145
post #59

Earlier quoted context omitted.

How many specialists did you go to before it was identified? How many other people with the condition were misidentified? I only say this because of a family member with a rare genetic condition. For years they were told it was something else, or told 'it was in their head'. The family member started a journal of their medical conditions and experiences that was detailed then brought that to their PC which whom then…

That all sounds shitty but I don't see how that's valuable information. They did eventually solve the problem and there's no comparison to some ML success story. Human minds can be really good at diagnostics and still fail sometimes when faced with very difficult cases. In my experience, ML would just classify everything as a very common disease and people would call it a success because it has an 80% effectiveness r…

>They did eventually solve the problem

One of the points taken should be that either via ML or human diagnostic is that these rare problems are either not diagnosed for long periods of time reducing quality of life or diagnosed posthumously.

The reduction of these measures are what we should use when making meat vs machine efficiency correlations.

Re: AI models that predict disease are not as accurate as reports might suggest

#146
post #49

Earlier quoted context omitted.

My partner had a clinician review her paperwork and say "why are you here" explaining the enhanced imaging was leading to tentative concerns being raised about structural change so small it was below the threshold for safe surgical treatment. Moral of the story: the imaging has got so good that diagnostics is now on the fringe of over diagnosing and the stats need to catch up

One of the first things AI will be really good at will be image post processing. Even then I, in case of medical diagnosis, I'd prefer to have an actual person compare the "RAW" image to whatever the AI came up with. Simply because post processing can create artifacts that can throw you of quite a bit. Regarding tue quality of imaging: I tend to agree, and the better imaging gets the more we will have to relly on hum…

The raw image is not good and isn’t usable. AI denoises and this is what makes it usable. Then it doubles the resolution.

There is no point in reviewing the raw image as it doesn’t add anything. A study is generally 300-1000 images. If you’re going to review the raw and therefore look at 600-2000 images you’ve just wasted everything AI gained and you might as well not use it.

I acquire the images and I look at what I get and re-run anything that’s got artifact. This is not unusual and generally happens due to movement, excessive image noise or incorrect parameter selection.

I don’t review the raw, but I do review the output. Keep in mind that MR images have always been very heavily processed at every stage of image formation as any way of squeezing more out the the acquisition will have time benefits.

There are a ton of ways a tech can create a misleading or faulty image with parameter selection which leads to an incorrect diagnosis. This could happen quite easily and AI being part of any error m is not something that keeps me awake at night, unlike other other parameters with which I have seen mistakes and have made them myself.

Re: AI models that predict disease are not as accurate as reports might suggest

#147
post #72
post #20

Technical (Honest) Solution: two holdouts 1. Involved in the build process 2. Never touched until paper metrics are being written, only run once Realistically, unlikely to occur however due to the incentives causing publication bias.

so the models that fail this one test never get published and the models that succeed get published. And all you have done is to publish a model that predicts that particular history, in other words data fitting.

> the models that fail this one test never get published

Not necessarily -- for comparison purposes one should include, and a negative result is not a bad outcome. But consider Edison's light bulb. Most of the failures don't matter, and a few might have interesting properties to reconsider or tweak down the line. But the major one folks care about is the one that worked.

> data fitting

Yes, models trained on data and not logic alone fit data.

Re: AI models that predict disease are not as accurate as reports might suggest

#148
I work in machine learning for digital Pathology and I think the big problem here is the divergence between publishing papers and real life helpful models. What we see is that in the literature you often see a model trained on data from a single lab and gets crazy good results. However, apply it to a different lab (not in the paper of course) and it sucks. So what we do in practice for our models is to train on many different labs at the same time and have a hold out test set with labs not covered during training. That way you get a robust model which works well in practice (but doesn't have the 99.9% metrics reported in papers). The second thing is looking at what task to let the machine do. We typically go for tasks which are boring / repetitive for the doctor, but visualize the result, so the doctor double checks it before making the diagnosis. Still saves a lot of time.

Re: AI models that predict disease are not as accurate as reports might suggest

#149
post #102

Earlier quoted context omitted.

> We humans are incredibly good at elimination of factors & differential diagnosis. I don't automatically buy this. Didn't heart attack care in the ER get dramatically better when people started following checklists? That suggests that human doctors aren't that great at even getting the basics correct. In addition, most doctors are below average . So, maybe the best doctors are better than the AI. However, I may not…

Using check lists means stabdardization, and that makes results compareable and reduces the risk of forgetting something under stress. Check lists have nothing to do with ML or AI so.

It wasn't just "forgetting". Every doctor had their own take on diagnosis and the checklist was actually better than a lot of them since the checklist was constructed from data.

Re: AI models that predict disease are not as accurate as reports might suggest

#150

I work in machine learning for digital Pathology and I think the big problem here is the divergence between publishing papers and real life helpful models. What we see is that in the literature you often see a model trained on data from a single lab and gets crazy good results. However, apply it to a different lab (not in the paper of course) and it sucks. So what we do in practice for our models is to train on many…

I also work in this field and what we typically do is a lot of nested cross-validations to get some bounds on a model building process and some idea of how it would perform on repeated unseen data. Data leakage is always on our mind and we do our best at all stages to avoid that. We also train on data from many sites. It can be done and it can be done properly. As you say, it is always best to collect some completely naive test set to back up the model-building process. If you design your pipeline properly, the test set should fall within the bounds you got during cross-validation. It all depends on how much data you have and I think as long as you design your pipelines with that in mind and acknowledge limitations with smaller datasets, then the research is valid and useful.
Post reply on HN