Live data from Hacker News

OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

theguardian.com

331–340 of 500 posts

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#331

Earlier quoted context omitted.

Doctors make errors all the time though, so the real argument is about the error percentage. If AIs is lower then it's safer (but it's hard to have that convo, I recognise). Besides; this article was about diagnosis not prescribing. It's pretty obvious, I think, that diagnosis is one area where AI will perform extremely well in the long run. I think there are two metrics; the first is outright misdiagnosis, which stu…

The bar for making ai useful is much lower though. It's enough to be better than nothing. Large populations also in the technically rich countries simply do not have access to a doctor. in Poland which has a free public Healthcare it takes literal years to get a single appointment sometimes.

Do you believe the issue is because they don't have enough technicians to diagnose or because they don't have enough x-ray machines? Or in a ER environment, how an AI would speed up things in a real way that improves patients' lives?

We just minted the term "cognitive debt" for software engineers that cannot keep up with what the AI spits out. How would that apply to ER doctors, or any other kind of doctor?

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#332
It would have been interesting to see how a doctor with access to LLMs would perform, compared to only LLMs and only doctors. If doctors with LLM access still score 67%, then someone with no medical knowledge could potentially score the same, which would make ER triage a replaceable task by AI. But I am sure that is not the case. Competent doctors with the background they have can use LLMs to brainstorm and analyze different paths and score higher.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#333
post #11

I'd be very very hesitant to trust studies like this. It's very easy to mess up these benchmarks. See for example this recent paper where AI managed to beat radiologists on interpreting x-rays... when the AI didn't even have access to the x-rays: https://arxiv.org/pdf/2603.21687 (on a pre existing "large scale visual question answering benchmark for generalist chest x-ray understanding" that wasn't intentionally mess…

I agree with you on this specific study, however, I can't really wrap my head about the fact that doctors will be better than AI models on the long-run. After all, medicine is all about knowledge, experience and intelligence (maybe "pattern recognition"), all those, we must assume that the best AI models (especially ones focusing solely in the medical field) would largely beat large majority of humans (aka doctors),…

You also have to assume advances in sensors and robotics (e.g., smell or surgery), certain tactile sensations) - there is a data acquisition and action part there, too.

In this study, I think there was an MD before the AI to enrich data.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#334

Earlier quoted context omitted.

I agree with you on this specific study, however, I can't really wrap my head about the fact that doctors will be better than AI models on the long-run. After all, medicine is all about knowledge, experience and intelligence (maybe "pattern recognition"), all those, we must assume that the best AI models (especially ones focusing solely in the medical field) would largely beat large majority of humans (aka doctors),…

> we must assume that the best AI models (especially ones focusing solely in the medical field) would largely beat large majority of humans (aka doctors), if we already have this assumption for software engineers You first have to assume this for software engineers. Not everyone agree with that (note: that doesn't mean the same people don't agree that AI is not _useful_). AIs still have a ton of issues that would be…

[dead]

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#335
post #211

Earlier quoted context omitted.

I used to think this too. But the past couple of years have soured my taste for "dismantle and replace" of vital institutions. I still think healthcare needs to be reformed, and I hope that insurance will someday be a thing of a past, but I've hung up my chain saw for now.

This is because "dismantle and replace" (or perhaps in other words, "defunding") is not a serious, viable solution to many of the societal issues we face. Things were ruined slowly. They unfortunately will need to be fixed very slowly too.

  > They unfortunately will need to be fixed very slowly too.
this can work until you hit a crisis point; i think one issue is we are sliding faster in the wrong direction (increasing bureaucracy, increasing fees, wait times, overwork etc) so "slowly" can work but only if its "fast enough" if you get what i mean (people are really suffering out there)

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#336
post #317

Off topic, is a “reject all and subscribe” cookie popup button legal? I thought websites have to make it as easy to give consent as withdraw consent[1] - and here one cannot withdraw consent without an extra step (subscribing). Instead I would expect access to the article, with same ads as in the “user consented” path, just not personalized. [1]: “The GDPR is specific that consent must be as 'easy to withdraw as to g…

No, typically it is not.

https://en.wikipedia.org/wiki/Consent_or_pay

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#338

Earlier quoted context omitted.

To answer your question: talking to a human. Medicine is so much more than "knowledge, experience, and pattern matching", as any patient ever can attest to. Why is it so hard for some people to understand that humans need other humans and human problems can't be solved with technology?

It seems likely to me that doctors whose job is almost or entirely about making diagnoses and prescribing treatments won't be able to keep up in the long run, where those who are more patient facing will still be around even after AI is better than us at just about everything. If I were picking a specialty now, I'd go with pediatrics or psychiatry over something like oncology.

You are confusing the job with a subset of tasks. Some tasks can be automated, some won't. That doesn't mean LLMs, which cannot tell how many r's are in strawberry, will replace anyone.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#339

Earlier quoted context omitted.

>What is the specific capability (or combination of capabilities) that people believe will remain permanently (or at least for decades) where a top medical AI cannot match or exceed the performance of a good human doctor? Let's put liability and ethics aside, let's be purely objective about it. You cannot simply put liability and ethics aside, after all there's Hippocatic oath that's fundamental to the practice physi…

"The boy who cried wolf" is a story about false positives, so if that's what you want to avoid then you want to get close to 100% specificity, and accept that there are many things that the tool will not catch. If, as you propose, the tool would mainly be used to create a low confidence list of potential problems that will be further reviewed by a human, then casting a wide net and calibrating for high sensitivity in…

The idea is to minimize the false positives "the boy who cried wolf" at the same time mitigate, or better eliminate false negatives. The main reason is that based on the physician in-the-loop, the system can be optimized for sensitivity but can be relaxed for specificity. Of course if can get both 100% sensitivity and specificity it will be great, but in life there's always a trade-off, c'est-la-vie.

In our novel ECG based CVD detection system we can get 100% sensitivity for both arrhythmia and ischemia, with inter-patient validation, not the biased intra-patient as commonly reported in literature even in some reputable conferences/journals. Specificity is still high around 90% but not yet 100% as in sensitivity but due to the physician-in-the-loop approach, which is a diagnostic requirement in the current practice of medicine, this should not be an issue.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#340
post #222

Earlier quoted context omitted.

To answer your question: talking to a human. Medicine is so much more than "knowledge, experience, and pattern matching", as any patient ever can attest to. Why is it so hard for some people to understand that humans need other humans and human problems can't be solved with technology?

I cannot wait until doctors are fully automated. Shouldn’t be long now, hopefully just a few years.

next year bro, I promise, now give me 60 billion more in funding
Post reply on HN