Live data from Hacker News

OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

theguardian.com

371–380 of 500 posts

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#371

> "An AI and a pair of human doctors were each given the same standard electronic health record to read" This is handicapping the human doctors abilities. There is a lot more information a human doctor can gather even with a brief observation of the patient.

They have covered this in the article. > But it is not curtains for emergency doctors yet, the researchers said. The study only tested humans against AIs looking at patient data that can be communicated via text. The AI’s reading of signals, such as the patient’s level of distress and their visual appearance, were not tested. That means the AI was performing more like a clinician producing a second opinion based on p…

> That means the AI was performing more like a clinician producing a second opinion based on paperwork.

That actually seems like a good application – automatically get a quick AI second opinion for everything; if it's dissenting the first/human medic can re-review, or comment why it's slop, or get a third/second-human opinion.

(I'm assuming most cases would be You're absolutely right, that's an astute diagnosis.)

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#372
post #142

Earlier quoted context omitted.

You’re holding on to the intuition (hope) that we are smarter than the LLMs in some hard to define way. Maybe. But it’s getting harder and harder to define a task that humans beat LLMs on. On pretty much any easily quantifiable test of knowledge or reasoning, the machines win. I agree experienced humans are still better on “judgement” tasks in their field. But the judgement tasks are kinda necessarily ones where ther…

> But it’s getting harder and harder to define a task that humans beat LLMs on. On pretty much any easily quantifiable test of knowledge or reasoning, the machines win. Quite to the contrary, I think it's extremely trivial to find a task where humans beat LLMs. For all the money that's been thrown at agentic coding, LLMs still produce substantially worse code than a senior dev. See my own prior comments on this for a…

> I doubt that LLMs will ever beat humans at this, but if LLMs can be proven to be good at point 2, then point 3 alone will not save human physicians.

Agree with your division but I'm baffled by this argument. If humans are better than machines at point 3 and can also use a machine to do point 2, then unless they have particularly terrible biases against taking point 2 data into account they're going to be strictly better than machines alone. Doctors have costs, but they're costs people/society are generally willing to underwrite, and misdiagnosis also has costs...

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#373

I’m in ophthalmology where AI diagnostics have been promised for almost a decade. We have FDA approved diagnostics for diabetic retinopathy screening that has been commercially available since 2018, and papers claiming board certified ophthalmologist level classification accuracy as far back as inceptionv3. Maybe it’s just an economic barrier but these tools still haven’t made any meaningful impact in the US. Other c…

AI diagnostics is maybe 60% the way there. Robotics is maybe 20% the way there. You'll have a job as a doctor for a good long while.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#374

Earlier quoted context omitted.

One doctor didn't want to give me ritalin, so i went to another one. One was against it, the other one saw it as a good idea. I would love to have real data, real statistics etc.

[flagged]

> Cool. Aren't LLMs already doing all the work that requires focus and intelligence instead of you?

So your solution is to outsource thinking and work? That'll work out great in the long run.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#375
post #10
post #5

Humans could not diagnose and treat me correctly . They almost killed me. Curious where I could feed my symptoms and the same data I gave to an ER to an AI to test it.

Chatgpt.com?

All the AI's are able to guess what is going on based on what information I gave the ER. I was under the impression that there is a different interface that does not redirect people to a real doctor and will try to act like a doctor which AI does not.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#376
post #220

Hold on. Does this mean ER diagnoses are marginally better than pure chance?

No, because randomly guessing from a list of diagnoses is not 50/50

And ER generally does not involve key decisions being made by someone isolated from the patient given only an incomplete set of notes to make their diagnosis

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#377
post #248
post #125

Earlier quoted context omitted.

> Diagnosis becomes cheaper & easier -> more time to actually talk to patients Unfortunately is this not likely to happen. More like: Diagnosis becomes cheaper & easier -> more patients a doctor is expected to see in the same period of time as before

What's unfortunate about that?

It is unfortunate because churning through patients quickly without actually listening to them well leads to worse outcomes

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#378
post #65

Earlier quoted context omitted.

5. Doctors delegate everything to AI assistants because humans are lazy, especially if those AI assistants are correct some significant portion of the time

Step 2 prevents that. It's not there by accident. They need to write down their (initial) diagnosis before the AI answer is shown.

Step 2 doesn't prevent it, because of step 4. AI becomes "upon further testing/examination/review we conclude that..."

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#379

Earlier quoted context omitted.

You are objectively incorrect. A fever is considered 100.4 or 38 C. Here are a few links to prove it: https://my.clevelandclinic.org/health/symptoms/10880-fever https://www.mayoclinic.org/diseases-conditions/fever/symptom... https://www.osfhealthcare.org/blog/whats-considered-a-fever-... https://www.brownhealth.org/be-well/fever-and-body-temperatu... https://www.childrensmercy.org/siteassets/media-documents-fo... I c…

Maybe you had trouble re-reading your own comment but I can tell by how you responded here (a cascade of links/references) and a snarky comment ("I can keep going if you'd like") that I'm sure the doctor was glad to be rid of you. You didn't say the doctor disputed you had a fever. You said the doctor told you the fever wasn't concern until 100.4. Which I'm guessing is your fault for misinterpreting. If you google ar…

> I clearly didn't have a fever

I actually did say that the doctor disputed I had a fever

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#380

Earlier quoted context omitted.

I agree with you on this specific study, however, I can't really wrap my head about the fact that doctors will be better than AI models on the long-run. After all, medicine is all about knowledge, experience and intelligence (maybe "pattern recognition"), all those, we must assume that the best AI models (especially ones focusing solely in the medical field) would largely beat large majority of humans (aka doctors),…

There are a few sides to medicine: 1) looking at tests and working out a set of actions 2) following a pathway based on diagnosis 3) pulling out patient history to work out what the fuck is wrong with someone. Once you have a diagnosis, in a lot of cases the treatment path is normally quite clear (ie patient comes in with abdomen pain, you distract the patient and press on their belly, when you release it they scream…

I'm a GP in the NHS - what is this DDx software that you talk about?
Post reply on HN