Live data from Hacker News

OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

theguardian.com

251–260 of 500 posts

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#251

Earlier quoted context omitted.

I would argue that the ED is the least similar to code. You have the most unknowns, unreliable data and history, non deterministic options and time constraints. An ER staff is frequently making inferences based on a variety of things like weather, what the pt is wearing, what smells are present, and a whole lot of other intangibles. Frequently the patients are just outright lying to the doctor. An AI will not pick up…

> An AI will not pick up on any of that. It will if it trains on data like that. It's all about the training data.

The user will be adversarial and probably learn new tricks to trick the machine, this is not solvable (only) via training data.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#252

Earlier quoted context omitted.

I would argue that the ED is the least similar to code. You have the most unknowns, unreliable data and history, non deterministic options and time constraints. An ER staff is frequently making inferences based on a variety of things like weather, what the pt is wearing, what smells are present, and a whole lot of other intangibles. Frequently the patients are just outright lying to the doctor. An AI will not pick up…

> An AI will not pick up on any of that. It will if it trains on data like that. It's all about the training data.

Unfortunately the training data is absolute garbage.

Diagnostic standards in (at least emergency, but I think other specialties) medicine are largely a joke -- ultimately it's often either autopsy or "expert consensus."

We get to bill more for more serious diagnoses. The amount of patients I see with a "stroke" or "heart attack" diagnosis that clearly had no such thing is truly wild.

We can be sued for tens of millions of dollars for missing a serious diagnosis, even if we know an alternative explanation is more likely.

If AI is able to beat an average doctor, it will be due to alleviating perverse incentives. But I can't imagine where we could get training data that would let it be any less of a fountain of garbage than many doctors.

Without a large amount of good training data, how could AI possibly be good at doctoring IRL?

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#253
post #11

I'd be very very hesitant to trust studies like this. It's very easy to mess up these benchmarks. See for example this recent paper where AI managed to beat radiologists on interpreting x-rays... when the AI didn't even have access to the x-rays: https://arxiv.org/pdf/2603.21687 (on a pre existing "large scale visual question answering benchmark for generalist chest x-ray understanding" that wasn't intentionally mess…

When you read through the article it shows that the gap between doctors and LLMs actually disappeared (in terms of statistical significance) once both were allowed to read the full case notes. The headline is quoting a number based on guessed diagnoses from nurse's notes. The LLM was happier to take guesses from the selected case studies than the doctors is my guess.

Not only is the study testing something which only vaguely resembles how doctors diagnose patients, but isolated accuracy percentages are also a terrible way to measure healthcare quality.

If 90% of patients have a cold, and 10% have metastatic aneuristic super-boneitis, then you can get 90% accuracy by saying every patient has a cold. I would expect a probabilistic token-prediction machine to be good at that. But hopefully, you can see why a human doctor might accept scoring a lower accuracy percentage, if it means they follow up with more tests that catch the 10% boneitis.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#254

I wouldn't put much weight in this study, but I think a lot of us can still attest to the usefulness of LLMs in self-diagnostics. The reality in the US is that it is difficult to get the attention and care of a doctor so we're left having to do it ourselves. 10 years ago you'd hear docs complaining about patients coming in with things they found on google but now I don't think there's an alternative. Case in point, I…

I don't think that using LLMs for medicine is an appropriate fix for the US's healthcare issues. Unless healthcare businesses decide to improve patient care with AI instead of increasing patients per day, I think it's going to make things even worse.

I'm not suggesting it as a fix. I'm saying it's the only option to get medical answers for many people.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#255

I wouldn't put much weight in this study, but I think a lot of us can still attest to the usefulness of LLMs in self-diagnostics. The reality in the US is that it is difficult to get the attention and care of a doctor so we're left having to do it ourselves. 10 years ago you'd hear docs complaining about patients coming in with things they found on google but now I don't think there's an alternative. Case in point, I…

I don't think that using LLMs for medicine is an appropriate fix for the US's healthcare issues. Unless healthcare businesses decide to improve patient care with AI instead of increasing patients per day, I think it's going to make things even worse.

Doctors using AI will probably just increasing the number of patients they see. But for me as patient AI is super useful to get a good handle on the situation before I see a doctor.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#256
post #38

Earlier quoted context omitted.

I agree with you on this specific study, however, I can't really wrap my head about the fact that doctors will be better than AI models on the long-run. After all, medicine is all about knowledge, experience and intelligence (maybe "pattern recognition"), all those, we must assume that the best AI models (especially ones focusing solely in the medical field) would largely beat large majority of humans (aka doctors),…

But liability and ethics cannot be put aside. If treatments were free of cost and perfectly address problems, then a correct diagnosis would always lead to the optimal patient outcome. In that scenario, AI diagnosis will be like code generation and go asymptotic to perfection as models improve. But a doctor's job in the real world today is to navigate a total mess of uncertainty: about the expected outcome of treatme…

Liability would put all this to bed. Is OpenAI liable for malpractice if it misdiagnoses your issue? No? Then it’s no substitute. Being right is not nearly as important as being responsible. Unfortunately, there is widespread perception that software defects are acceptable, whereas operating on the wrong leg isn’t.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#257
post #243

Earlier quoted context omitted.

If you prefer an LLM to a human doctor, you deserve an LLM instead of a human doctor, and I wish you get it.

I would use one for sure. Much of medicine is getting tests / labs booked fighting to get certain medicines. Doctors will barely give you 5 minutes only deal with one issue per visit, rarely are available and going into an office can make you sicker. An llm with Doctor powers could offer more. I don't think we are at the surgery point but we are past getting notes and medicine's refilled.

So why not order your own labs? I'm sure you can think of ways to get your own medications if you are sufficiently convinced that this is the best course of action for your health.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#258
post #95

Earlier quoted context omitted.

To answer your question: talking to a human. Medicine is so much more than "knowledge, experience, and pattern matching", as any patient ever can attest to. Why is it so hard for some people to understand that humans need other humans and human problems can't be solved with technology?

>Medicine is so much more than "knowledge, experience, and pattern matching", as any patient ever can attest to. Humans (doctors/nurses) can still be there to make you feel the warmth of humanity in your darkest times, but if a machine is going to perform better at diagnosing (or perhaps someday performing surgery), then I want the machine. Even now, I'll take a surgeon that's a complete jerk over a nice surgeon any…

> Even now, I'll take a surgeon that's a complete jerk over a nice surgeon any day, because if they've got that job even as a jerk they've got to be good at their jobs.

This seems like an incredibly poor line of reasoning.

Hospitals are often desperate for surgeons. The poorly mannered ones are often deeply unsatisfied, angry at the grueling lives they've opted into, and the hospitals can't replace them. The market is not exactly at work here.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#259
post #179

Earlier quoted context omitted.

To answer your question: talking to a human. Medicine is so much more than "knowledge, experience, and pattern matching", as any patient ever can attest to. Why is it so hard for some people to understand that humans need other humans and human problems can't be solved with technology?

LLMs are a distillation of human.

Human language that is.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#260

Earlier quoted context omitted.

medical industry must be going for some long term achievement in how much they disbelieve, mistreat, and degrade women going to them. I wonder how many units of their training courses are spent on this and how much is spent on the cultural reinforcement of it.

Yes, let's pretend that the bias does not exist, that is helpful. It certainly doesn't have to do with the fact that it's currently a 60/40 split in active male vs female physicians. Or that women are more likely to be taken seriously by doctors: * https://www.health.harvard.edu/pain/the-dangerous-dismissal-of-womens-pain * https://pmc.ncbi.nlm.nih.gov/articles/PMC10937548/ Are you really unwilling to admit that such…

This seems like an especially bad faith interpretation of the comment you were responding to.
Post reply on HN