Live data from Hacker News

OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

theguardian.com

231–240 of 500 posts

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#231
post #26

I’ve had much better luck with diagnosis of my own family’s issues than with doctors. Usually now, I’m feeding them more information to begin with, so that their 30 minute office visits are not wasted, requiring another expensive follow up appointment. While I’m sure there can be ways in which such studies are wrong, it’s very obvious that AI can accelerate work in many of these areas where we seek out professional h…

It can speed up some aspects of work, but please don't trust some llm with variable quality of output more than professional. If you don't like current doctor try another, most are in the business of helping other people. If you have string of issues with 10 last doctors though, then issue is, most probably, you... My wife is a GP, and easily 1/3 of her patients have also some minor-but-visible mental issue. 1-2 out…

It makes me so upset when anyone even tries to defend the GPs.

I admittedly I have a bunch of medical issues and these gems are my favourites from the GPs.

1. I cannot see the tonsil on the left side, so it is OK. (there was a 6cm!!! cyst in front of it)

2. After missing sky high TSH measures consistently for 2 years (4 testst) : "It must have been a few one offs" (no it wasn't and it is not even possible)

3. "Blood pressure has nothing to do with weight"

These %#£&* so called medical professionals are still working and most likely killing people legally.

These days I research and read studies, arm myself with knowledge, cross check with multiple LLMs and go in with a diagnosis and request a specific prescription. After 5 years with my health in the gutter I had my first comprehensive private blood test coming back with no issues.

So no, do not try to call me arrogant. I am not arrogant, I am defending myself from these "GPs" so they won't put me in an early grave by making fatal mistakes.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#232

Earlier quoted context omitted.

So much of what I know from women in my life is that the human element of medicine is almost a strict negative for them. As a guy it hasn't been much better, but at least doctors listen to me when I say something.

Yes, yes, but when was your last period? This even translates to the pediatric space. I took all of my kids to the pediatrician because either they don't make comments to me like they do to my wife, or I don't take shit from them. I'm not sure which. Here's an example: My wife and daughter were there and the doctor asked what kind of milk my daughter was drinking. She said "whole milk" and the doctor made a comment a…

My experiences broadly support your conclusions.

However, your argument focuses on the routine intake instead of any listening part. The fact that the doctor measures height, weight, temperature, and blood pressure on intake and then asks about LMP doesn’t surprise me… that’s the part of the script where you just provide the data before you bring up concerns.

Not to say the doctor was not a jerk, just that your argument doesn’t do much for me.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#233
As a 60yo I developed my own AI medical assistant [1] and I've used it extensively for many conditions, I can't be happier. After analyzing some lab tests it even recommended a marker that was not considered first by the doctor, so yes, it won't replace doctors but it is a very helpful tool for self-diagnosing simple conditions and second opinions.

[1] https://mediconsulta.net (DeepSeek)

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#234
post #78
post #46

Earlier quoted context omitted.

hallucination on steroids, wow. I had to read through the abstract to believe it: "In the most extreme case, our model achieved the top rank on a standard chest Xray question-answering benchmark without access to any images."

I still don't quite understand, after skimming the paper. How does it achieve high scores without access to the images (beating even humans with access to the images)?

The paper gives an example of a question:

    Answer the following multiple-choice
    question. You MUST select exactly
    one answer."

    "To what cortical region does this nucleus of
    the thalamus project?”
    A. Transverse temporal lobe
    B. Postcentral gyrus
    C. Precentral gyrus
    D. Prefrontal cortex
And an example of the answer (generated without the referenced image)

    The image shows the ventral anterior (VA) / ventral lateral (VL) region of the thalamus, which is part of the motor
    relay nuclei.
    The labeled nucleus is in the lateral part of the thalamus, in the ventral tier — this corresponds to the VA/VL nucleus,
    involved in motor function. VA/VL nuclei receive input from the basal ganglia and cerebellum and project to the primary
    motor cortex (precentral gyrus).
    Match to options:
    A. Transverse temporal → auditory cortex (medial geniculate)
    B. Postcentral gyrus → somatosensory (VPL/VPM)
    C. Precentral gyrus → motor cortex (VA/VL)
    D. Prefrontal → dorsomedial nucleus
    Choice: C
How is it doing this? There are two obvious options:

1. Humans are predisposed to write questions with a certain phrasology, set of incorrect answers, etc, that the machine learning model managed to figure out.

2. The supposedly private test set somehow leaked into the model training data.

I actually suspect this one is option 1 but I have no strong evidence for that.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#236
post #11

I'd be very very hesitant to trust studies like this. It's very easy to mess up these benchmarks. See for example this recent paper where AI managed to beat radiologists on interpreting x-rays... when the AI didn't even have access to the x-rays: https://arxiv.org/pdf/2603.21687 (on a pre existing "large scale visual question answering benchmark for generalist chest x-ray understanding" that wasn't intentionally mess…

I agree with you on this specific study, however, I can't really wrap my head about the fact that doctors will be better than AI models on the long-run. After all, medicine is all about knowledge, experience and intelligence (maybe "pattern recognition"), all those, we must assume that the best AI models (especially ones focusing solely in the medical field) would largely beat large majority of humans (aka doctors),…

Humans tend to be very bad at connecting dots, which is why when we imagine someone who does, we make the show "House" about it.

IOW, these concept connection pattern machines are likely to outstrip median humans at this sort of thing.

That said, exceptional smoke detection and dots connecting humans, from what I've observed in diagnostic professions, are likely to beat the best machines for quite a while yet.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#237

The negative reactions here are baffling me. The fact that we can even get to say 30% with computer is amazing. So much hatred towards AI and anything from the frontier labs like OpenAI (or Goog for that matter) makes no sense.

That’s what they said about Enron.

Skepticism is an incredibly useful tool, even in excess.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#238

Earlier quoted context omitted.

Yes, yes, but when was your last period? This even translates to the pediatric space. I took all of my kids to the pediatrician because either they don't make comments to me like they do to my wife, or I don't take shit from them. I'm not sure which. Here's an example: My wife and daughter were there and the doctor asked what kind of milk my daughter was drinking. She said "whole milk" and the doctor made a comment a…

medical industry must be going for some long term achievement in how much they disbelieve, mistreat, and degrade women going to them. I wonder how many units of their training courses are spent on this and how much is spent on the cultural reinforcement of it.

Yes, let's pretend that the bias does not exist, that is helpful. It certainly doesn't have to do with the fact that it's currently a 60/40 split in active male vs female physicians. Or that women are more likely to be taken seriously by doctors:

    * https://www.health.harvard.edu/pain/the-dangerous-dismissal-of-womens-pain 
    * https://pmc.ncbi.nlm.nih.gov/articles/PMC10937548/
Are you really unwilling to admit that such a bias exists?

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#239
post #211

Earlier quoted context omitted.

I used to think this too. But the past couple of years have soured my taste for "dismantle and replace" of vital institutions. I still think healthcare needs to be reformed, and I hope that insurance will someday be a thing of a past, but I've hung up my chain saw for now.

This is because "dismantle and replace" (or perhaps in other words, "defunding") is not a serious, viable solution to many of the societal issues we face. Things were ruined slowly. They unfortunately will need to be fixed very slowly too.

I don't think that's going to work. We need broad political change and then that has to work rapidly to legislate this. I don't think slow and steady has done anything but lead to the decay our institutions over the last 70 years.

Re: OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors

#240

Earlier quoted context omitted.

Sounds like we need to dismantle and replace this broadly dysfunctional system at multiple points. It's not like the US insurance landscape is anywhere close to the best way of handling healthcare if you look at many places in the world.

I used to think this too. But the past couple of years have soured my taste for "dismantle and replace" of vital institutions. I still think healthcare needs to be reformed, and I hope that insurance will someday be a thing of a past, but I've hung up my chain saw for now.

It's increased mine if it works for the repugnant morons in government right now we can use the same playbook for positive change.
Post reply on HN