Live data from Hacker News

OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

theguardian.com

321–330 of 500 posts

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#321

Earlier quoted context omitted.

One doctor didn't want to give me ritalin, so i went to another one. One was against it, the other one saw it as a good idea. I would love to have real data, real statistics etc.

[flagged]

Because i actually have real ADHD.

I have it so strong, that after I was preparing myself, my work desc, my books everything, i was starring into the books i wanted to learn for 15-30 minutes unable to just start or do anything.

With ritalin, i might have this mental block to, but its overcome in a few seconds.

I went from a 'nearly/borderline failing grade' to the nearly the best grade in just one year.

This changed significantly were I am today.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#322

Earlier quoted context omitted.

Code is pretty much the perfect use case for LLMs… text-based, very pattern-oriented, extremely limited complexity compared to biological systems, etc. I suspect even prose is largely considered acceptable in professional uses because we haven’t developed a sensitivity to the artifice, and we probably won’t catch up to the LLMs in that arms race for a bit. However, we always manage to develop a distaste for cheap imi…

And with the code, the closer you come to the physical world the worse LLMs fair. Claude can’t really write Openscad and when I was debugging some map projections code last week it struggled a lot more than usual.

Until anthropic hire or steal code from acquired companies and train with it.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#323

Earlier quoted context omitted.

> After all, medicine is all about knowledge, experience and intelligence So is... everything? LLMs are really really good at knowledge. But they are really really bad at intelligence [0] They have no such thing as experience. Do not fool yourself, intelligence and knowledge are not the same thing. It is extremely easy to conflate the two and we're extremely biased to because the two typically strongly correlate. But…

Yeah, I mean, I don't know where all of this is going, but I do think that the ancients cared WAY more about "embodied knowledge" than we do, and I suspect we're about to find out a lot more about what that is and why it matters.

There's a lot of definitions of bodies. Though I'm unconvinced one is needed. A brain in a box is capable of interacting with its environment far more than such a thing could even a decade ago. Is it the body or the interaction?

As we advance we always need to answer more nuanced questions. You're right that the nature of progress is... well... progress

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#324

Earlier quoted context omitted.

> we must assume that the best AI models (especially ones focusing solely in the medical field) would largely beat large majority of humans (aka doctors), if we already have this assumption for software engineers You first have to assume this for software engineers. Not everyone agree with that (note: that doesn't mean the same people don't agree that AI is not _useful_). AIs still have a ton of issues that would be…

Doctors make errors all the time though, so the real argument is about the error percentage. If AIs is lower then it's safer (but it's hard to have that convo, I recognise). Besides; this article was about diagnosis not prescribing. It's pretty obvious, I think, that diagnosis is one area where AI will perform extremely well in the long run. I think there are two metrics; the first is outright misdiagnosis, which stu…

The bar for making ai useful is much lower though. It's enough to be better than nothing.

Large populations also in the technically rich countries simply do not have access to a doctor.

in Poland which has a free public Healthcare it takes literal years to get a single appointment sometimes.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#325

> "An AI and a pair of human doctors were each given the same standard electronic health record to read" This is handicapping the human doctors abilities. There is a lot more information a human doctor can gather even with a brief observation of the patient.

You could say the same about the Ai. Ai is incredibly well suited for extracting knowledge through chats. In this regard. A doctor also just have 15 minutes for an interview. An Ai can be with the patient for days leading up to a consultation. So if we remove this "handicap" this Ai will likely really start to win.

It’s the ER. People aren’t always in a position to “chat” when they go there.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#326

> "An AI and a pair of human doctors were each given the same standard electronic health record to read" This is handicapping the human doctors abilities. There is a lot more information a human doctor can gather even with a brief observation of the patient.

You could say the same about the Ai. Ai is incredibly well suited for extracting knowledge through chats. In this regard. A doctor also just have 15 minutes for an interview. An Ai can be with the patient for days leading up to a consultation. So if we remove this "handicap" this Ai will likely really start to win.

Chat seems like a really bad way to get patient information. You'll miss out on various cues doctors will use to diagnose you. People can get ashamed of their symptoms and may try to hide them.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#327

Earlier quoted context omitted.

I agree with you on this specific study, however, I can't really wrap my head about the fact that doctors will be better than AI models on the long-run. After all, medicine is all about knowledge, experience and intelligence (maybe "pattern recognition"), all those, we must assume that the best AI models (especially ones focusing solely in the medical field) would largely beat large majority of humans (aka doctors),…

> we must assume that the best AI models (especially ones focusing solely in the medical field) would largely beat large majority of humans (aka doctors), if we already have this assumption for software engineers You first have to assume this for software engineers. Not everyone agree with that (note: that doesn't mean the same people don't agree that AI is not _useful_). AIs still have a ton of issues that would be…

In some subfields, like detection of security weaknesses in obscure C code, AI is already better than software engineers.

It is capable of sifting through enormous reams of data without ever zoning out etc. Once patients routinely use various wearables etc., they, too, will produce heaps of data to be analyzed, and AI will be the thing to go to when it comes to anomaly detection.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#328
Not long ago I started having an issue with my eye. I called around and they said I should get seen ASAP, same day if possible, but it wasn’t worth the ER and it was a five day wait for an appointment.

I was pretty freaked out. During that time, I tried diagnosing it with AI. When I finally got to the appointment, the actual doctor sat down, looked at all the unremarkable images, asked me one (1) question, ordered another image and diagnosed the issue. When I looked back, in all that time, the AI had mentioned it exactly one time early on, ruled it out immediately based on a flawed understanding of the symptoms, and never brought it up again.

Just my anecdotal evidence, but I’d never trust any AI on its own. My doctor can use it if they want, I can’t.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#329

Earlier quoted context omitted.

I agree with you on this specific study, however, I can't really wrap my head about the fact that doctors will be better than AI models on the long-run. After all, medicine is all about knowledge, experience and intelligence (maybe "pattern recognition"), all those, we must assume that the best AI models (especially ones focusing solely in the medical field) would largely beat large majority of humans (aka doctors),…

>What is the specific capability (or combination of capabilities) that people believe will remain permanently (or at least for decades) where a top medical AI cannot match or exceed the performance of a good human doctor? Let's put liability and ethics aside, let's be purely objective about it. You cannot simply put liability and ethics aside, after all there's Hippocatic oath that's fundamental to the practice physi…

"The boy who cried wolf" is a story about false positives, so if that's what you want to avoid then you want to get close to 100% specificity, and accept that there are many things that the tool will not catch. If, as you propose, the tool would mainly be used to create a low confidence list of potential problems that will be further reviewed by a human, then casting a wide net and calibrating for high sensitivity instead does make sense.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#330
post #11

I'd be very very hesitant to trust studies like this. It's very easy to mess up these benchmarks. See for example this recent paper where AI managed to beat radiologists on interpreting x-rays... when the AI didn't even have access to the x-rays: https://arxiv.org/pdf/2603.21687 (on a pre existing "large scale visual question answering benchmark for generalist chest x-ray understanding" that wasn't intentionally mess…

Yup, there's a reason while ROC is a thing in data science. You can build a 99% accurate cancer detector that's just a slip of paper saying 'you don't have cancer', but everybody understands its worthless intuitively. With more complex setups, that intuition goes away.
Post reply on HN