I have generally moved from bearish to bullish on the future of current AI technology, but the continued inaccuracy with basic facts all while the models significantly improve continues to give me significant pause. As an example, creating recipes with Claude Opus based on flavor profiles and preferences feels magical, right up until the point at which it can't accurately convert between tablespoons and teaspoons. It…
If Claude occasionally overestimates the conversion, that might be an artifact of Australian tablespoons being different (4tsp/20mL vs 3tsp/15mL in the US). This error could at least be explained as a complication of the real world. (If it's saying 3.14tsp or 2tsp then I have no idea)
Ontario auditors find doctors' AI note takers routinely blow basic facts
131–140 of 141 posts
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#132Earlier quoted context omitted.
Hey, I'm agreement with you. I meant that these limited extracts do need to be recorded, that's all. Read the rest of the comment :)
Oops... I am deeply sorry, thank you for the heads up! It seems I've myself committed a cardinal sin that I am usually quick to point in others - rushing to reply without comprehending the full message. (Meta-oops: I realized how LLM-ish it sounds. Quick, reboot before my cover is blown!) I happen to believe that the flaw being discussed IS fundamental and inherent in the design and architecture of LLM - this is why…
Which both of us do.
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#133Earlier quoted context omitted.
> Real intelligence means you have to say "I don't know" when you don't know I have met many supposedly intelligent, certainly high status, humans who don't appear to be able to do that either. I have more confidence we can train AIs to do it, honestly.
While it is true that there are people who do not admit they are wrong when they factually are, your assertion glosses over the fact that most of the people we maintain in our social circle are people we trust through our experiences with them to be honest.
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#134I have generally moved from bearish to bullish on the future of current AI technology, but the continued inaccuracy with basic facts all while the models significantly improve continues to give me significant pause. As an example, creating recipes with Claude Opus based on flavor profiles and preferences feels magical, right up until the point at which it can't accurately convert between tablespoons and teaspoons. It…
When I see how LLMs are capable of essentially prompt and context engineering for themselves, it makes me think they won't need human guidance forever.
When it comes to simple fact-based tasks that have a concrete methodology, it is no surprise to me that LLMs aren't the right tool, and I believe it's a failure of the harness to not recognize those types of tasks and handle them with a more concretely functioning tool instead of relying on statistical probabilities in the LLM "brain" to spit out the correct number to a math problem.
In the same sense that LLMs can use "skills" when necessary, it should have tools or possibly even specialized "brains" for it to pass of certain types of tasks to.
I'm starting to feel that our first form of AGI is not going to be a single brain but an elaborate system of harnesses, multiple LLM models, skills, domain and task specialized subsystems it passes tasks off to, etc. Whether we get there with current LLM technology before some other evolution in AI is the question, to me.
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#135Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#136Can someone who is a more AI heavy user explain what is going on? I would expect an "AI Note Taker" to faithfully transcribe the entire conversation. With the same quality I see in a lot of automated video subtitles.. ie they use the wrong word a lot but it's easy to tell what they mean by context. Are these tools instead immediately summarising the whole thing, and that summary is the artifact? Because that is a bey…
Obviously this results in hallucinations, mistaken implications, & inaccurately assumed context.
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#137Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#138The linked report seems almost useless -- it doesn't say anything about an error rate or a sample size, so it's a mystery whether 9 out of 20 systems “fabricated information and made suggestions to patients' treatment plans” ten out of ten times, or one out of a thousand times. If we just postulate that the systems have a high error rate, I wonder why they are being adopted. They seem extremely easy to test, so I don…
>If we just postulate that the systems have a high error rate, I wonder why they are being adopted. From the article: "While 30 percent of a platform’s evaluation score depended solely on whether they had a domestic presence in Ontario, the accuracy of medical notes contributed only 4 percent to the total score." Accuracy wasn't really part of the scoring, Ontario doesn't care about it.
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#139I have generally moved from bearish to bullish on the future of current AI technology, but the continued inaccuracy with basic facts all while the models significantly improve continues to give me significant pause. As an example, creating recipes with Claude Opus based on flavor profiles and preferences feels magical, right up until the point at which it can't accurately convert between tablespoons and teaspoons. It…
I was skeptical that LLMs could be the right path to AGI, but then I kept being impressed by how much further we could take it by expanding upon the way we use it, the harnesses we use with LLMs, and better context engineering. When I see how LLMs are capable of essentially prompt and context engineering for themselves, it makes me think they won't need human guidance forever. When it comes to simple fact-based tasks…
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#140Earlier quoted context omitted.
> It is pretty clear that initial accuracy issues will become less and less of a problem as these technologies mature. What do you base this on? As someone who can both see the amazing things genAI can do, and who sees how utterly flawed most genAI output is, it's not obvious to me. I'm working with Claude every day, Opus 4.7, and reviewing a steady stream of PRs from coworkers who are all-in, not just using due to c…
> That is a vintage hallucination that could've come right out of GPT 2.0. That's because, despite the many claims to the contrary, the models haven't actually gotten any smarter. They are still just token prediction engines at the end of the day, without any understanding of what they are doing. That's why one should not rely on them.
I'm not sure, but it seems to me that if scale or small architectural tweaks were going to solve comprehension, they would have done so by now.