Live data from Hacker News

OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

theguardian.com

381–390 of 500 posts

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#382

Earlier quoted context omitted.

You think current ER people work in complete silence? No words uttered?

You think that they have “days leading up to consultation”? Please don’t be so disingenuous; I’m sure you know exactly what the person you’re replying to meant.

> I’m sure you know exactly what the person you’re replying to meant.

No.

There are a lot of different modus operandi, and you can always find an outlier.

> Please don’t be so disingenuous;

Ditto

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#383
post #77

radiology already had its "AI beats doctors" moment. radiologists are still here. what changed first was the workflow, not the specialty. er is probably next.

I don't think radiology has had that moment at all. Computer programming is much closer, if not, at that moment right now.

no programming it's still just tool use for CRUD applications with react and tailwind

complex systems programming is just so unreliable and foolish to use LLMs to do anything important

companies adopting it for more safety critical systems are just already seeing the problems pile on and we're seeing news about it almost every day on Hacker News

If the tool can make something look smart but isn't necessarily correct, lazy employed humans will just defer to it, especially when their lazy greedy bosses tell them to, and everybody loses over time (except the stakeholders that just jump companies anyway after they made their money)

It's just sad to see these really unwise and inexperienced sentiments repeated ad nauseam

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#384

Earlier quoted context omitted.

I agree with you on this specific study, however, I can't really wrap my head about the fact that doctors will be better than AI models on the long-run. After all, medicine is all about knowledge, experience and intelligence (maybe "pattern recognition"), all those, we must assume that the best AI models (especially ones focusing solely in the medical field) would largely beat large majority of humans (aka doctors),…

It's having a general understanding/view of the "baseline", aka healthy anatomy. This is something LLMs will never have, that's why never have true reasoning, for the lack of "worldview" and they never know if they are hallucinating. To aid doctors, we don't need LLMs but rather, computer vision, pattern recognition as you correctly point out. But it's important not to rely on it. Doctors can easily recognize and cor…

>It's having a general understanding/view of the "baseline", aka healthy anatomy. This is something LLMs will never have

You're making the mistake of conflating AI with LLMs.

I don't think LLMs will reliably be better than a board of doctors. But an Expert System probably will (if it isn't already). That's literally what they were created for.

The biggest downside of LLMs IMO isn't the millions of Jules wasted on training models that are ultimately used to create funny images of cats with lasers. It's that all that money isn't being invested into truly helpful AI systems that will actually improve and save our lives, such as medical expert systems.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#385

Earlier quoted context omitted.

It's having a general understanding/view of the "baseline", aka healthy anatomy. This is something LLMs will never have, that's why never have true reasoning, for the lack of "worldview" and they never know if they are hallucinating. To aid doctors, we don't need LLMs but rather, computer vision, pattern recognition as you correctly point out. But it's important not to rely on it. Doctors can easily recognize and cor…

>It's having a general understanding/view of the "baseline", aka healthy anatomy. This is something LLMs will never have You're making the mistake of conflating AI with LLMs. I don't think LLMs will reliably be better than a board of doctors. But an Expert System probably will (if it isn't already). That's literally what they were created for. The biggest downside of LLMs IMO isn't the millions of Jules wasted on tra…

I am quite surprised that expert systems are not already used in this area (and others). As you say, this is exactly what they are meant for.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#386

> "An AI and a pair of human doctors were each given the same standard electronic health record to read" This is handicapping the human doctors abilities. There is a lot more information a human doctor can gather even with a brief observation of the patient.

They have covered this in the article. > But it is not curtains for emergency doctors yet, the researchers said. The study only tested humans against AIs looking at patient data that can be communicated via text. The AI’s reading of signals, such as the patient’s level of distress and their visual appearance, were not tested. That means the AI was performing more like a clinician producing a second opinion based on p…

> The study only tested humans against AIs looking at patient data that can be communicated via text.

This is like saying that LLMs can evaluate paintings better than art experts. But only when looking at data that can be communicated via text.

Of course they can, because it makes no sense to do such a thing.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#387
post #131
post #49

I'm surprised at both the article and the paper - both seem very hyperbolic. This is LLMs competing against doctors in a way that is heavily weighted in the LLMs favour, which does not represent clinical practice. These reasoning cases are not benchmarks for doctors, they are learning tools. I think it's important to note that diagnosis also relies on accurate description of the patient in the first place, and the in…

Also, you need to see an analysis of the incorrect calls. The goal of a human Dr is not to get the highest accuracy, it's to limit total harm to the patient. There can be cases where the odds favor picking X (but it may not be by that much), but the safe thing to do is to rule out some other option first, or start a safe treatment that covers several other possible options. Simply getting the "high score" on this eva…

Yeah 100% this. We've all used AI. It's obvious that it can sometimes outperform humans in a "did it get the right answer" benchmark while being wildly worse overall because of worse failure modes.

I bet the AI's incorrect answers are less "I don't know, let's get a second opinion" and more "you're perfectly fine, 0% chance this is cancer".

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#388
post #11

I'd be very very hesitant to trust studies like this. It's very easy to mess up these benchmarks. See for example this recent paper where AI managed to beat radiologists on interpreting x-rays... when the AI didn't even have access to the x-rays: https://arxiv.org/pdf/2603.21687 (on a pre existing "large scale visual question answering benchmark for generalist chest x-ray understanding" that wasn't intentionally mess…

Ultimatly you'd want humans and AI to study separately cases separately and independtly, and flag cases that have been found by only one analysis so that a separate analysis is done by a second pair of eyes.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#389

Earlier quoted context omitted.

You've witnessed a dismantle and replace effort by the right wing that wishes to squeeze everything to make rich people more money. An effort by the left would destroy the private insurance scheme and build up medicare. Completely different and you'd get something functional. When the wrong targets get destroyed, everyone suffers. When parasitic forces are destroyed, the system functions better. It's the difference b…

We already had an effort by the left. You can “no true scotsmen” if you want, but it represents the reality of what will happen when ideals clash a sector that makes up 18% of the GDP. What’s going to be different now than in 2010?

Are you referring to the ACA here? That was a compromise bill that props up the current system in the US, primarily created by right leaning centrists.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#390

Earlier quoted context omitted.

It's having a general understanding/view of the "baseline", aka healthy anatomy. This is something LLMs will never have, that's why never have true reasoning, for the lack of "worldview" and they never know if they are hallucinating. To aid doctors, we don't need LLMs but rather, computer vision, pattern recognition as you correctly point out. But it's important not to rely on it. Doctors can easily recognize and cor…

>It's having a general understanding/view of the "baseline", aka healthy anatomy. This is something LLMs will never have You're making the mistake of conflating AI with LLMs. I don't think LLMs will reliably be better than a board of doctors. But an Expert System probably will (if it isn't already). That's literally what they were created for. The biggest downside of LLMs IMO isn't the millions of Jules wasted on tra…

The nature of expert systems is to become experts on a system.

The reason you need a doctor, or more often, let's be honest, a good nurse, is because systems can fail in any one of 10000 as yet undiscovered ways. New nurses. New residents. New techs. And on and on and on. All the measurements you're feeding to the system are an amalgamation of the potential errors of a potentially different set of professionals each time you move a patient through the enterprise.

Full disclosure, my first startup was building PACS and RTP software back before AI reading was a thing. Current startup working across dental and medical. Rethinking the link between oral and systemic health. Partner has been in the C-suite of several hospitals over the past few decades and now runs large healthcare delivery networks.

The reason you can't hand things over to AI, is precisely because there are so many humans in the system. Each of whom are fallible. Human experts are quicker to catch it. Expert systems are not. At least not any ES or AI I've seen. And I've been going to, for instance, RSNA, for well over 25 years.

If you have an ES or AI in the system, you would naturally put the same professionals responsible for catching human screwups, in charge of catching AI and ES screw ups. Even if these AI's turn 100% accurate based on the inputs they are given, that professional would still be responsible for catching those bad inputs.

Example, it's never happened to one of my companies knock on wood, but I have seen cases of radiation therapy patients being incorrectly dosed. The doctor almost never was the one who miffed in the situation, but ultimately, s/he's responsible.

Why? Bad input should have been caught.

Another example, situations where you operate on the wrong side of the body because someone prepped the wrong leg. Surgeon didn't do the prep. Whoever did do the prep may have simply relied on the software. But the software was wrong. May have been anything. Point is, the team is good, but everyone just fell into too complacent of a pattern with each other and their tools.

Trust is good. Complacency is not.

The same will hold true for AI team members that integrate into these environments. It's just another "team member", and it better have a "monitor". If not, you're asking for trouble.

The "monitor" ultimately responsible for everything will continue to be the provider. Any change in that reality will take decades. (And in the end, they probably will not change the current system in that regard.)

Post reply on HN