OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
381–390 of 500 posts
Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
#382Earlier quoted context omitted.
You think current ER people work in complete silence? No words uttered?
You think that they have “days leading up to consultation”? Please don’t be so disingenuous; I’m sure you know exactly what the person you’re replying to meant.
No.
There are a lot of different modus operandi, and you can always find an outlier.
> Please don’t be so disingenuous;
Ditto
Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
#383radiology already had its "AI beats doctors" moment. radiologists are still here. what changed first was the workflow, not the specialty. er is probably next.
I don't think radiology has had that moment at all. Computer programming is much closer, if not, at that moment right now.
complex systems programming is just so unreliable and foolish to use LLMs to do anything important
companies adopting it for more safety critical systems are just already seeing the problems pile on and we're seeing news about it almost every day on Hacker News
If the tool can make something look smart but isn't necessarily correct, lazy employed humans will just defer to it, especially when their lazy greedy bosses tell them to, and everybody loses over time (except the stakeholders that just jump companies anyway after they made their money)
It's just sad to see these really unwise and inexperienced sentiments repeated ad nauseam
Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
#384Earlier quoted context omitted.
I agree with you on this specific study, however, I can't really wrap my head about the fact that doctors will be better than AI models on the long-run. After all, medicine is all about knowledge, experience and intelligence (maybe "pattern recognition"), all those, we must assume that the best AI models (especially ones focusing solely in the medical field) would largely beat large majority of humans (aka doctors),…
It's having a general understanding/view of the "baseline", aka healthy anatomy. This is something LLMs will never have, that's why never have true reasoning, for the lack of "worldview" and they never know if they are hallucinating. To aid doctors, we don't need LLMs but rather, computer vision, pattern recognition as you correctly point out. But it's important not to rely on it. Doctors can easily recognize and cor…
You're making the mistake of conflating AI with LLMs.
I don't think LLMs will reliably be better than a board of doctors. But an Expert System probably will (if it isn't already). That's literally what they were created for.
The biggest downside of LLMs IMO isn't the millions of Jules wasted on training models that are ultimately used to create funny images of cats with lasers. It's that all that money isn't being invested into truly helpful AI systems that will actually improve and save our lives, such as medical expert systems.
Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
#385Earlier quoted context omitted.
It's having a general understanding/view of the "baseline", aka healthy anatomy. This is something LLMs will never have, that's why never have true reasoning, for the lack of "worldview" and they never know if they are hallucinating. To aid doctors, we don't need LLMs but rather, computer vision, pattern recognition as you correctly point out. But it's important not to rely on it. Doctors can easily recognize and cor…
>It's having a general understanding/view of the "baseline", aka healthy anatomy. This is something LLMs will never have You're making the mistake of conflating AI with LLMs. I don't think LLMs will reliably be better than a board of doctors. But an Expert System probably will (if it isn't already). That's literally what they were created for. The biggest downside of LLMs IMO isn't the millions of Jules wasted on tra…
Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
#386> "An AI and a pair of human doctors were each given the same standard electronic health record to read" This is handicapping the human doctors abilities. There is a lot more information a human doctor can gather even with a brief observation of the patient.
They have covered this in the article. > But it is not curtains for emergency doctors yet, the researchers said. The study only tested humans against AIs looking at patient data that can be communicated via text. The AI’s reading of signals, such as the patient’s level of distress and their visual appearance, were not tested. That means the AI was performing more like a clinician producing a second opinion based on p…
This is like saying that LLMs can evaluate paintings better than art experts. But only when looking at data that can be communicated via text.
Of course they can, because it makes no sense to do such a thing.
Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
#387I'm surprised at both the article and the paper - both seem very hyperbolic. This is LLMs competing against doctors in a way that is heavily weighted in the LLMs favour, which does not represent clinical practice. These reasoning cases are not benchmarks for doctors, they are learning tools. I think it's important to note that diagnosis also relies on accurate description of the patient in the first place, and the in…
Also, you need to see an analysis of the incorrect calls. The goal of a human Dr is not to get the highest accuracy, it's to limit total harm to the patient. There can be cases where the odds favor picking X (but it may not be by that much), but the safe thing to do is to rule out some other option first, or start a safe treatment that covers several other possible options. Simply getting the "high score" on this eva…
I bet the AI's incorrect answers are less "I don't know, let's get a second opinion" and more "you're perfectly fine, 0% chance this is cancer".
Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
#388I'd be very very hesitant to trust studies like this. It's very easy to mess up these benchmarks. See for example this recent paper where AI managed to beat radiologists on interpreting x-rays... when the AI didn't even have access to the x-rays: https://arxiv.org/pdf/2603.21687 (on a pre existing "large scale visual question answering benchmark for generalist chest x-ray understanding" that wasn't intentionally mess…
Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
#389Earlier quoted context omitted.
You've witnessed a dismantle and replace effort by the right wing that wishes to squeeze everything to make rich people more money. An effort by the left would destroy the private insurance scheme and build up medicare. Completely different and you'd get something functional. When the wrong targets get destroyed, everyone suffers. When parasitic forces are destroyed, the system functions better. It's the difference b…
We already had an effort by the left. You can “no true scotsmen” if you want, but it represents the reality of what will happen when ideals clash a sector that makes up 18% of the GDP. What’s going to be different now than in 2010?
Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
#390Earlier quoted context omitted.
It's having a general understanding/view of the "baseline", aka healthy anatomy. This is something LLMs will never have, that's why never have true reasoning, for the lack of "worldview" and they never know if they are hallucinating. To aid doctors, we don't need LLMs but rather, computer vision, pattern recognition as you correctly point out. But it's important not to rely on it. Doctors can easily recognize and cor…
>It's having a general understanding/view of the "baseline", aka healthy anatomy. This is something LLMs will never have You're making the mistake of conflating AI with LLMs. I don't think LLMs will reliably be better than a board of doctors. But an Expert System probably will (if it isn't already). That's literally what they were created for. The biggest downside of LLMs IMO isn't the millions of Jules wasted on tra…
The reason you need a doctor, or more often, let's be honest, a good nurse, is because systems can fail in any one of 10000 as yet undiscovered ways. New nurses. New residents. New techs. And on and on and on. All the measurements you're feeding to the system are an amalgamation of the potential errors of a potentially different set of professionals each time you move a patient through the enterprise.
Full disclosure, my first startup was building PACS and RTP software back before AI reading was a thing. Current startup working across dental and medical. Rethinking the link between oral and systemic health. Partner has been in the C-suite of several hospitals over the past few decades and now runs large healthcare delivery networks.
The reason you can't hand things over to AI, is precisely because there are so many humans in the system. Each of whom are fallible. Human experts are quicker to catch it. Expert systems are not. At least not any ES or AI I've seen. And I've been going to, for instance, RSNA, for well over 25 years.
If you have an ES or AI in the system, you would naturally put the same professionals responsible for catching human screwups, in charge of catching AI and ES screw ups. Even if these AI's turn 100% accurate based on the inputs they are given, that professional would still be responsible for catching those bad inputs.
Example, it's never happened to one of my companies knock on wood, but I have seen cases of radiation therapy patients being incorrectly dosed. The doctor almost never was the one who miffed in the situation, but ultimately, s/he's responsible.
Why? Bad input should have been caught.
Another example, situations where you operate on the wrong side of the body because someone prepped the wrong leg. Surgeon didn't do the prep. Whoever did do the prep may have simply relied on the software. But the software was wrong. May have been anything. Point is, the team is good, but everyone just fell into too complacent of a pattern with each other and their tools.
Trust is good. Complacency is not.
The same will hold true for AI team members that integrate into these environments. It's just another "team member", and it better have a "monitor". If not, you're asking for trouble.
The "monitor" ultimately responsible for everything will continue to be the provider. Any change in that reality will take decades. (And in the end, they probably will not change the current system in that regard.)