LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes
11–17 of 17 posts
Re: LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes
#12The whole thing is nasty partly because it isn't just AI problem. When checking the human labels that we used in our evals 40 out of 142 answer keys claiming absence were wrong. Tricky one.
Re: LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes
#13Not sure about the paper but the results make sense, we see it in PDF extraction too. Fields that aren't in the document are being made up 11 to 40% of the time depending on the API. The whole thing is nasty partly because it isn't just AI problem. When checking the human labels that we used in our evals 40 out of 142 answer keys claiming absence were wrong. Tricky one.
Do you have an example?
I'm in healthcare and ambient documentation is obviously a huge thing now but I don't have any experience with it. We have anywhere from 5-10 companies reach out a week trying to sell us on their product and the demos are mostly okay (though you can tell they're rely on the happy path through a lot of it), but we haven't actually pulled the trigger on anything. Thanks in advance.
Re: LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes
#14People using AI like this should be run out of their jobs.
Re: LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes
#15Not sure about the paper but the results make sense, we see it in PDF extraction too. Fields that aren't in the document are being made up 11 to 40% of the time depending on the API. The whole thing is nasty partly because it isn't just AI problem. When checking the human labels that we used in our evals 40 out of 142 answer keys claiming absence were wrong. Tricky one.
> Fields that aren't in the document are being made up 11 to 40% of the time depending on the API. Do you have an example? I'm in healthcare and ambient documentation is obviously a huge thing now but I don't have any experience with it. We have anywhere from 5-10 companies reach out a week trying to sell us on their product and the demos are mostly okay (though you can tell they're rely on the happy path through a l…
Receipts where there's no subtotal line, only a total. One model returns a subtotal anyway, "63.000", "88.000", numbers it copied or computed from elsewhere in the document. Ours invented a 2% discount on a receipt that has no discount line.
TV ad contracts where the station's address isn't on the contract. Several models fill it with a real address that is on the page, just the agency's or the advertiser's. Given your industry this pattern should worry you most, since these aren't made up from nothing. Just the model tripping and taking it from the wrong place, which might look plausible on the surface.
Here's what I'd suggest for your vendor demos. Run them on a handful of your own documents where you know a field is genuinely absent, and count how many come back filled. Their own samples won't tell you that.
All the raw outputs are here if you want to look: https://velrim.com/research/fabrication-on-absent-fields
Re: LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes
#16Not sure about the paper but the results make sense, we see it in PDF extraction too. Fields that aren't in the document are being made up 11 to 40% of the time depending on the API. The whole thing is nasty partly because it isn't just AI problem. When checking the human labels that we used in our evals 40 out of 142 answer keys claiming absence were wrong. Tricky one.
> Fields that aren't in the document are being made up 11 to 40% of the time depending on the API. Do you have an example? I'm in healthcare and ambient documentation is obviously a huge thing now but I don't have any experience with it. We have anywhere from 5-10 companies reach out a week trying to sell us on their product and the demos are mostly okay (though you can tell they're rely on the happy path through a l…
TV ad contracts where the station's address isn't on the contract. Several models fill it with a real address that is on the page, just the agency's or the advertiser's. Given your industry this pattern should worry you most, since these aren't made up from nothing. Just the model tripping and taking it from the wrong place, which might look plausible on the surface.
Here's what I'd suggest for your vendor demos. Run them on a handful of your own documents where you know a field is genuinely absent, and count how many come back filled. Their own samples won't tell you that.
All the raw outputs are public, so you can check the examples above yourself. My reply to you got filtered so I'm omitting the link. If you want to read the writeup, search for "velrim" and "fabrication on absent fields" for the article plus the repo.
Re: LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes
#17Earlier quoted context omitted.
I understand the answer to the question I'm about to ask, but how does any human read that abstract and not think, "This is entirely too many colons." EDIT: And Pangram agrees that the abstract is 100% AI generated.
While I agree I have to say that AI-detectors are complete bullshit