KPMG pulls report on AI usage due to apparent hallucinations
11–20 of 37 posts
Re: KPMG pulls report on AI usage due to apparent hallucinations
#12I guess nobody ever got fired for paying KPMG and friends for an expensive report that supported their priors.
Re: KPMG pulls report on AI usage due to apparent hallucinations
#13The crazy thing is the level of effort to say, "have a sub agent validate all references and figures" is so low. I'm paraphrasing, but you don't need much more than that. It would have prevented 99% of the face palms. I use this regularly for my personal financial research system. Even flagship models make mistakes. Though currently the issue is usually the model using a figure from and older report. Cross-check redu…
dont be so sure they didnt. they can go back and forth hallucinating with each other
Then have another set of agents, with skills like web browsing (to verify that links actually exist, maybe that references and abstracts actually match, etc), have one engineer (or agent) write a small script to help with this (just make sure you test it, and a bit).
So your work is not verified until your references table is 90% green checkmarks, maybe with uncertainty figures.
A human can then verify the ones with under 90% certainty.
This alone gets you a long way there. Does not costs the millions they're being paid.
It's quite interesting that these companies marketed themselves as them best of the best in excellence, accept no mistakes. I can imagine the countless keynotes and books about this. Or the sales pitches.
Has always been a lie, they just understood how to hide it. Today they don't, and it's embarrassing.
Re: KPMG pulls report on AI usage due to apparent hallucinations
#14The crazy thing is the level of effort to say, "have a sub agent validate all references and figures" is so low. I'm paraphrasing, but you don't need much more than that. It would have prevented 99% of the face palms. I use this regularly for my personal financial research system. Even flagship models make mistakes. Though currently the issue is usually the model using a figure from and older report. Cross-check redu…
Re: KPMG pulls report on AI usage due to apparent hallucinations
#15Earlier quoted context omitted.
dont be so sure they didnt. they can go back and forth hallucinating with each other
This is where the absolutism of let agents to 100% of the work fails. You get adversarial agents pulling all reverences into a table, they might miss some, so run this a few times. Then have another set of agents, with skills like web browsing (to verify that links actually exist, maybe that references and abstracts actually match, etc), have one engineer (or agent) write a small script to help with this (just make s…
How about the author actually reads the finished report a couple of times and checks all the references?
It really is the lowest bar - even lower maybe than running a spell check.
Re: KPMG pulls report on AI usage due to apparent hallucinations
#16Re: KPMG pulls report on AI usage due to apparent hallucinations
#17Re: KPMG pulls report on AI usage due to apparent hallucinations
#18Earlier quoted context omitted.
This is where the absolutism of let agents to 100% of the work fails. You get adversarial agents pulling all reverences into a table, they might miss some, so run this a few times. Then have another set of agents, with skills like web browsing (to verify that links actually exist, maybe that references and abstracts actually match, etc), have one engineer (or agent) write a small script to help with this (just make s…
> A human can then verify the ones with under 90% certainty. How about the author actually reads the finished report a couple of times and checks all the references? It really is the lowest bar - even lower maybe than running a spell check.
But then you wouldn't be embracing the new agentic ways of working!
Re: KPMG pulls report on AI usage due to apparent hallucinations
#19Earlier quoted context omitted.
dont be so sure they didnt. they can go back and forth hallucinating with each other
This is where the absolutism of let agents to 100% of the work fails. You get adversarial agents pulling all reverences into a table, they might miss some, so run this a few times. Then have another set of agents, with skills like web browsing (to verify that links actually exist, maybe that references and abstracts actually match, etc), have one engineer (or agent) write a small script to help with this (just make s…