OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
411–420 of 500 posts
Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
#412Earlier quoted context omitted.
To answer your question: talking to a human. Medicine is so much more than "knowledge, experience, and pattern matching", as any patient ever can attest to. Why is it so hard for some people to understand that humans need other humans and human problems can't be solved with technology?
If you read the study, the whole conclusion is much less spectacular than the article. What the article really pushes happened: patients -> AI -> diagnosis (you know, with a camera, or perhaps a telephone I guess) What REALLY happened patients -> nurse/MD -> text description of symptoms -> MD -> question (as in MD asked a relevant diagnostic question, such as "is this the result of a lung infection?", or "what lab te…
100% of the cases where some headline makes big claims about "AI" based on some study, you take a good hard look at the study and none of the big claims stand on their own.
It's all heavily spinned, taken out of context, editorialized... It's become almost a hobby of mine lately. And I am glad for have read so many papers and reasoned critically about methods and statistics. But it is also scary to realize just how much people take at face value of bombastic interpretations of datasets that support no such claim or much weaker versions only.
Chasing down sources is something that I often do and I've learned that people take a lot of liberty when divulging opinions about sources they don't think will be checked. Even in high trust environments. I have first hand received work by post-doctoral fellows where some articles in the bibliography didn't even exist.
Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
#413I know a cardiologist who founded a training & knowledge base startup for doctors. He once told me (that was before LLMs), that it’s super common to tell a patient that the doc needs to look up sthg in their patient history, to then instead google the symptoms. Or, even more often, quickly text a colleague. I have no way of knowing if this is true. But I‘d rather had a complete, guided prompt be the basis of a diagno…
> quickly text a colleague. This is still common and useful to gut check and make sure you aren't missing something. Source: wife is a doctor.
Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
#414Fifty percent accuracy. That's terrible.
Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
#415Earlier quoted context omitted.
Not only is the study testing something which only vaguely resembles how doctors diagnose patients, but isolated accuracy percentages are also a terrible way to measure healthcare quality. If 90% of patients have a cold, and 10% have metastatic aneuristic super-boneitis, then you can get 90% accuracy by saying every patient has a cold. I would expect a probabilistic token-prediction machine to be good at that. But ho…
What percentage of patients have blood clots in their lungs and a history of lupus, like the article described? That's not on the same level as a common cold at all .
> In one case in the Harvard study, a patient presented with a blood clot to the lungs and worsening symptoms.
That's a single anecdotal fluke from the study, which is misleadingly used to represent the headlining percentages.
If you read the linked paper, it says the LLMs did not outperform any group of doctors in the most important cases:
> The median proportion of cannot-miss diagnoses included for o1-preview was 0.92 [interquartile range (IQR) 0.62 to 1.0], although this was not significantly higher than GPT-4, attending physicians, or residents.
And again, the bigger issue is that skimming nurse's notes and predicting the next tokens, as the study made the doctors do, is not how doctors diagnose medical conditions.
Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
#416Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
#417Besides for myself and wife, I've also used LLMs to diagnose my dogs. Convinced there's a huge opportunity for AI based veterinary, especially one which then performs bidding across the local veterinary clinics to perform the care/surgeries. I've noticed that local vets vary in price by more than an order of magnitude. My 80 year old mother and mother inlaw have been regularly scammed by over charging vets, and with…
Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
#418I'd be very very hesitant to trust studies like this. It's very easy to mess up these benchmarks. See for example this recent paper where AI managed to beat radiologists on interpreting x-rays... when the AI didn't even have access to the x-rays: https://arxiv.org/pdf/2603.21687 (on a pre existing "large scale visual question answering benchmark for generalist chest x-ray understanding" that wasn't intentionally mess…
I agree with you on this specific study, however, I can't really wrap my head about the fact that doctors will be better than AI models on the long-run. After all, medicine is all about knowledge, experience and intelligence (maybe "pattern recognition"), all those, we must assume that the best AI models (especially ones focusing solely in the medical field) would largely beat large majority of humans (aka doctors),…
Do we have that assumption? I don't think there's a consensus on it yet, just various camps of people proselytizing the other camps based on how much or little they use AI.
Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
#419Earlier quoted context omitted.
I would use one for sure. Much of medicine is getting tests / labs booked fighting to get certain medicines. Doctors will barely give you 5 minutes only deal with one issue per visit, rarely are available and going into an office can make you sicker. An llm with Doctor powers could offer more. I don't think we are at the surgery point but we are past getting notes and medicine's refilled.
So why not order your own labs? I'm sure you can think of ways to get your own medications if you are sufficiently convinced that this is the best course of action for your health.
Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
#420Earlier quoted context omitted.
What percentage of patients have blood clots in their lungs and a history of lupus, like the article described? That's not on the same level as a common cold at all .
> One experiment focused on 76 patients who arrived at the emergency room of a Boston hospital. > In one case in the Harvard study, a patient presented with a blood clot to the lungs and worsening symptoms. That's a single anecdotal fluke from the study, which is misleadingly used to represent the headlining percentages. If you read the linked paper, it says the LLMs did not outperform any group of doctors in the mos…