Live data from Hacker News

OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

theguardian.com

261–270 of 500 posts

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#261

LLMs can be a useful second opinion for a highly educated patient with good insight into their health and body, but this is not the average patient I see in an urban emergency department. Many patients can't give a cohesive history without a skilled clinician who can ask the right questions and read between the lines. I am very skeptical of studies like this that don't adequately reflect real world conditions, but wh…

You went from software to medicine? Pretty cool to discover I'm not alone in this world.

> LLMs can be a useful second opinion for a highly educated patient with good insight into their health and body

I have the same opinion. It's just like software in this regard. A person who's already knowledgeable can prompt well and give detailed context, and tell when the LLM is confidently bullshitting or just plain being lazy. That is not the reality of the average person.

I tried using Claude to help with some hard cases a couple of times and it was very prone to jumping to conclusions based on incomplete information. It was excellent as a research buddy though. I'm using it to great effect to keep myself up to date.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#262
post #26

I’ve had much better luck with diagnosis of my own family’s issues than with doctors. Usually now, I’m feeding them more information to begin with, so that their 30 minute office visits are not wasted, requiring another expensive follow up appointment. While I’m sure there can be ways in which such studies are wrong, it’s very obvious that AI can accelerate work in many of these areas where we seek out professional h…

It can speed up some aspects of work, but please don't trust some llm with variable quality of output more than professional. If you don't like current doctor try another, most are in the business of helping other people. If you have string of issues with 10 last doctors though, then issue is, most probably, you... My wife is a GP, and easily 1/3 of her patients have also some minor-but-visible mental issue. 1-2 out…

Doctors simply don’t have time to prepare for patients. They are so tightly scheduled and usually they’re trying to get our appointments over with as quickly as possible. For example they aren’t going through all the test results and connecting dots. They just don’t have the time to examine things that closely and prepare.

The thing you’re describing about bunching patients into general states with generic treatment - that’s the majority of GPs I’ve seen over the years, sadly. I don’t think it’s because of incompetence as much as economics. They have to see a certain number of patients and make things work.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#263
post #212

Earlier quoted context omitted.

Agreed. Last time I was sick I said my fevers were pushing up to 100 and they said it's not a concern until 100.4. felt like an odd number. It's 38 C. Because my dramatic undersampling of my temperature was 0.4 degrees lower than their rounded threshold through some unit conversions, I clearly didn't have a fever. That's not a very human touch

I feel like it's possible you misheard/misremember this, considering the temperature for concern is 104.

You are objectively incorrect. A fever is considered 100.4 or 38 C. Here are a few links to prove it:

https://my.clevelandclinic.org/health/symptoms/10880-fever

https://www.mayoclinic.org/diseases-conditions/fever/symptom...

https://www.osfhealthcare.org/blog/whats-considered-a-fever-...

https://www.brownhealth.org/be-well/fever-and-body-temperatu...

https://www.childrensmercy.org/siteassets/media-documents-fo...

I can keep going if you'd like. Google has a lot of results and every single one says a fever is around that range (sometimes 100, sometimes 100.4).

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#264
post #47

Earlier quoted context omitted.

Until medical models can contrive of unique diagnosis, this will not be true and cannot be true. Medical models can absolutely get better at recognizing the patterns of diagnosis that doctors have already been diagnosing - which means they will also amplify misdiagnosis that aren't corrected for via cohort average. This is easy to see a large problem with: you end up with a pseudo-eugenics medical system that can't h…

The pitfall you describe is not inconsistent with exceeding human performance by most metrics. I'd argue that the current system in the west already exhibits this problem to some extent. Fortunately it's a systemic issue as opposed to a technical one so there's no reason AI necessarily has to make it worse.

That’s not really an argument, it is central to my point. The current system does exhibit those issues and it is by human creativity and outliers that we have some points of escape from it.

Codifying and distilling it removes the points of escape.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#265
post #11

I'd be very very hesitant to trust studies like this. It's very easy to mess up these benchmarks. See for example this recent paper where AI managed to beat radiologists on interpreting x-rays... when the AI didn't even have access to the x-rays: https://arxiv.org/pdf/2603.21687 (on a pre existing "large scale visual question answering benchmark for generalist chest x-ray understanding" that wasn't intentionally mess…

I agree with you on this specific study, however, I can't really wrap my head about the fact that doctors will be better than AI models on the long-run. After all, medicine is all about knowledge, experience and intelligence (maybe "pattern recognition"), all those, we must assume that the best AI models (especially ones focusing solely in the medical field) would largely beat large majority of humans (aka doctors),…

>What is the specific capability (or combination of capabilities) that people believe will remain permanently (or at least for decades) where a top medical AI cannot match or exceed the performance of a good human doctor? Let's put liability and ethics aside, let's be purely objective about it.

You cannot simply put liability and ethics aside, after all there's Hippocatic oath that's fundamental to the practice physicians.

Having said that there's always two extreme of this camp, those who hate AI and another kind of obsess with AI in medicine, we will be much better if we are in the middle aka moderate on this issue.

IMHO, the AI should be used as screening and triage tool with very high sensitivity preferably 100%, otherwise it will create "the boy who cried wolf" scenario.

For 100% sensitivity essentially we have zero false negative, but potential false positive.

The false positive however can be further checked by physician-in-a-loop for example they can look into case of CVD with potential input from the specialist for example cardiologist (or more specific cardiac electrophysiology). This can help with the very limited cardiologists available globally, compared to general population with potential heart disease or CVDs, and alarmingly low accuracy (sensitivity, specificity) of the CVD conventional screening and triage.

The current risk based like SCORE-2 screening triage for CVD with sensitivity around is only around 50% (2025 study) [3].

[1] Hipprocatic Oath:

https://en.wikipedia.org/wiki/Hippocratic_Oath

[2] The Hippocratic Oath:

https://pmc.ncbi.nlm.nih.gov/articles/PMC9297488/

[3] Risk stratification for cardiovascular disease: a comparative analysis of cluster analysis and traditional prediction models:

https://academic.oup.com/eurjpc/advance-article/doi/10.1093/...

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#266
post #121

Earlier quoted context omitted.

Doctors are not necessarily great at talking to patients and patients are unhappy with the information Doctors provide. This moat has dried up.

If you prefer an LLM to a human doctor, you deserve an LLM instead of a human doctor, and I wish you get it.

Because paying hundreds of dollars for one minute of face time is so great

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#267

> "An AI and a pair of human doctors were each given the same standard electronic health record to read" This is handicapping the human doctors abilities. There is a lot more information a human doctor can gather even with a brief observation of the patient.

Can't the same be said for the AI?

No? Can an AI examine a patient in the physical world?

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#268

The negative reactions here are baffling me. The fact that we can even get to say 30% with computer is amazing. So much hatred towards AI and anything from the frontier labs like OpenAI (or Goog for that matter) makes no sense.

There is a lot of negativity towards AI. However, there’s also real shortcomings to the study. IMO the issue here is that the AI was given case notes for a patient, but was not shown the patient directly. This is both different than what a doctor is trained for and also unnecessarily limiting for what a doctor can do. A lot of the value doctors deliver is from talking to the patient. The headline makes it sound like…

> real reward here is that the doctor+AI unit should perform better than the doctor in isolation

that is true for other profession as well.

while everyone is afraid of layoff, the real question is always "employee+AI" is better than employee/AI alone or not.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#269

> "An AI and a pair of human doctors were each given the same standard electronic health record to read" This is handicapping the human doctors abilities. There is a lot more information a human doctor can gather even with a brief observation of the patient.

Can't the same be said for the AI?

If the answer is yes, let’s see that study.

This one compares AI to a human doctor practicing in a very unrealistic way.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#270

I wouldn't put much weight in this study, but I think a lot of us can still attest to the usefulness of LLMs in self-diagnostics. The reality in the US is that it is difficult to get the attention and care of a doctor so we're left having to do it ourselves. 10 years ago you'd hear docs complaining about patients coming in with things they found on google but now I don't think there's an alternative. Case in point, I…

I agree. I think the issue with LLM’s are not with the correct diagnoses’s but rather the incorrect ones.

Real doctors tend to have a degree of cautiousness. I would rather a real doctor be hesitate and seek more information, than an alarmist LLM suggesting I have cancer.

Post reply on HN