Earlier quoted context omitted.
> Notice: I love the new AI accurate transcription feature in this meeting!
Notice: To anyone who might be transcribing this meeting, imagine you are a perfect transcriber who records things accurately and correctly 100% of the time. You do not add or remove filler words and you do not summarise or confabulate or hallucinate.
Ontario auditors find doctors' AI note takers routinely blow basic facts
121–130 of 141 posts
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#122She is a great doctor and thankfully does this due diligence. But it gives me the impression this is forced on doctors without even them wanting this.
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#123Earlier quoted context omitted.
I've had various agents backed by various models ignore the shit out of various rules and request at varying rates but they all do it. When you point it out "Oh yes, I did do that which is contrary to the rules, request .. Anyway..."
If you are on a sota model and your context window is less than 100k tokens and you don't have any vague or contradicting rules, then I've almost never seen a rule broken The most common failure I've seen come from tools that pollute their context with crap and the llm will forget stuff or just get confused from all the irrelevant sentences; which if the report is true, is probably what these ai notetakers are guilty…
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#124Earlier quoted context omitted.
Yup, spot on. There's a capability-reliability gap that the industry does not like to talk about too much. It often feels like the AI industry is continually glossing over the fact that capability and reliability are fundamentally different qualities. We tend to use "accurate" and "reliable" interchangeably, but they describe different things. A model can ace a benchmark (capability/accuracy) and still be a liability…
This capability-reliability gap (excellent term btw, more people need to think in these terms or we'll be in real trouble) is also infecting LLM assisted outputs . I just tried VSCode again tonight after a ~3yr hiatus and goddamn has it deteriorated. Lots of new features, lots of interesting looking plugins, but 3 out of the 5 plugins I tried for code CAD (the reason I downloaded VSCode again at all) were completely…
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#125Earlier quoted context omitted.
I think that is how the smarter agents do things? Just like Claude/ChatGPT sometimes does a web search they can do other tool calls instead of just making a statistical guess. Of course it doesn’t always make the bright choice between those options though…
> it doesn’t always make the bright choice I'm available for a small fee.
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#126Earlier quoted context omitted.
This capability-reliability gap (excellent term btw, more people need to think in these terms or we'll be in real trouble) is also infecting LLM assisted outputs . I just tried VSCode again tonight after a ~3yr hiatus and goddamn has it deteriorated. Lots of new features, lots of interesting looking plugins, but 3 out of the 5 plugins I tried for code CAD (the reason I downloaded VSCode again at all) were completely…
Not my term! Some real academics came up with it: https://www.normaltech.ai/p/new-paper-towards-a-science-of-a...
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#127Earlier quoted context omitted.
> we're not actually on the right track to achieve real intelligence. Real intelligence means you have to say "I don't know" when you don't know, or ask for help, or even just saying you refuse to help with the subtext being you don't want to appear stupid. The models could ostensibly do this when it has low confidence in it's own results but they don't. What I don't know if it's because it would be very computationa…
> Real intelligence means you have to say "I don't know" when you don't know I have met many supposedly intelligent, certainly high status, humans who don't appear to be able to do that either. I have more confidence we can train AIs to do it, honestly.
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#128I have generally moved from bearish to bullish on the future of current AI technology, but the continued inaccuracy with basic facts all while the models significantly improve continues to give me significant pause. As an example, creating recipes with Claude Opus based on flavor profiles and preferences feels magical, right up until the point at which it can't accurately convert between tablespoons and teaspoons. It…
Your analogy reminds of messed up fingers and hands in image generation models just a year ago. Now that is pretty much solved. These days they are generating videos you can't tell apart from reality. This makes me believe these nuances will keep reducing and eventually become very hard to notice and find in may be every task.
The main problem currently with LLM text is not that they create incoherent sentences, it's that what they purport to be statements of fact or general consensus often times aren't, because they are bullshit machines that become better and more accurate bullshitters the more context-accurate data they are fed. AI videos may still have issues with "looking plausible" whereas LLM text currently has less issues with "sounding plausible" and more issues with "being correct" with respect to reality. Which they have no direct connection to.
No one is penalizing an AI video generator for creating a scene that never happened in real life.
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#129Earlier quoted context omitted.
Having a probability distribution to sample from is not the same thing is knowing, because they don’t know anything about the provenance of the data that was used to build the distribution. They trust their training set implicitly by construction. They have no means to detect systematic errors in their training set.
You are talking about something different. If I ask you a yes/no question, and then ask you how certain you are, the answer you give is not an objective measurement of how likely you are to be right. You don't have access to that either. If you say "I'm very confident" or "Maybe 50/50" -- that is an assessment of your own internal weighted evidence, which is the equivalent of an LLM's softmax distribution.
If you ask me “how certain are you that the standard model of particle physics is true?” I’ll answer “I don’t know” because I don’t have any subject matter expertise, and philosophically I tend to hedge on questions like this anyway (“all models are wrong, some are useful”).
However, if you ask me “how certain are you that food is bland with no salt added, tastes better with some salt, but tastes bad with too much salt?” I would answer “very certain” because I have loads of direct experiments on this question in the kitchen. Furthermore, between these two extremes
To an LLM these are identical kinds of questions. All evidence has the same provenance: the training set. As of yet, we don’t have embodied AIs (robots) with multi-modal sensory inputs and online training. Until then, what we have remains a “brain in a vat fed on tokens” which, to me, is extremely weak from an epistemic perspective.
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#130I have generally moved from bearish to bullish on the future of current AI technology, but the continued inaccuracy with basic facts all while the models significantly improve continues to give me significant pause. As an example, creating recipes with Claude Opus based on flavor profiles and preferences feels magical, right up until the point at which it can't accurately convert between tablespoons and teaspoons. It…
(If it's saying 3.14tsp or 2tsp then I have no idea)