Live data from Hacker News

EY Canada published a cybersecurity report and most citations were hallucinated

gptzero.me

111–120 of 156 posts

Re: EY Canada published a cybersecurity report and most citations were hallucinated

#111
post #8

The problem we're seeing across many professions is AI output is not getting vetted by knowledgeable people, whether it's an experienced analyst, senior engineer, expert attorney, or the resident physician. At best they skim, at worst they don't even see it at all before it's published, pushed to production, distributed to clients, or submitted to the court. In many cases the skills are available in house to do the n…

As an attorney, I feel like vetting AI output takes longer than just doing it from scratch, let alone versus just using a traditional form. With AI, I have to read through everything, often explain why it's wrong, and then rewrite everything anyways. I mean, I get way more billables, but I think it's symptomatic of how AI loses its advantage of being quick and accessible to those who don't understand the subject matt…

You can also feed the document or source file to another frontier-level model, ideally two others, and tell it to vet it aggressively. The goal is to goad the models into erring on the side of false positive findings rather than potentially missing true positives.

I find that if Gemini Pro agrees with Claude Opus 4.8 and GPT 5.5 on something, it's almost certainly correct at a level where I wouldn't be likely to catch any errors myself.

Re: EY Canada published a cybersecurity report and most citations were hallucinated

#112
post #8

The problem we're seeing across many professions is AI output is not getting vetted by knowledgeable people, whether it's an experienced analyst, senior engineer, expert attorney, or the resident physician. At best they skim, at worst they don't even see it at all before it's published, pushed to production, distributed to clients, or submitted to the court. In many cases the skills are available in house to do the n…

As an attorney, I feel like vetting AI output takes longer than just doing it from scratch, let alone versus just using a traditional form. With AI, I have to read through everything, often explain why it's wrong, and then rewrite everything anyways. I mean, I get way more billables, but I think it's symptomatic of how AI loses its advantage of being quick and accessible to those who don't understand the subject matt…

Another attorney here. I understand your plight. But I can't believe law firms are sending out briefs and opinions without carefully checking all of the citations. I mean, even when Lexis or Westlaw identifies an (actual) case on point, you still have to check if the case has been overturned, whether it is truly on point, or if it can be distinuished from your case. So even if the cited case is not a halucination, someone would still have to read and analyze the cited case in the context of the present case.

Re: EY Canada published a cybersecurity report and most citations were hallucinated

#113
post #40

Is there any source with just the plain text? The css styling is headache inducing and reader mode doesn't work or has been defeated.

Firefox has a handy "Reader view" (Opt + CMD + R on Mac) that you can activate to get a stripped down view of just the text on the page. Unfortunately, it also removes the images which contain some of the sources they use.

Re: EY Canada published a cybersecurity report and most citations were hallucinated

#114

I don't quite get it why they can't take another LLM and vet the output of the first with the second one. Surely they would not have the same hallucinations and would be able to detect hallucinations of the earlier LLM. Maybe it would cost too much in terms of tokens? I don't know but I would expect it to be realtively easy for an LLM to detect "hallucinations".

I am not exactly sure if this would solve the overall problem. The main one being lack of oversight. The solution to a social issue generally isn’t to throw more technology at it.

Re: EY Canada published a cybersecurity report and most citations were hallucinated

#115
post #8

The problem we're seeing across many professions is AI output is not getting vetted by knowledgeable people, whether it's an experienced analyst, senior engineer, expert attorney, or the resident physician. At best they skim, at worst they don't even see it at all before it's published, pushed to production, distributed to clients, or submitted to the court. In many cases the skills are available in house to do the n…

>>The problem we're seeing across many professions is AI output is not getting vetted by knowledgeable people

I am particularly interested in Education and Human Knowledge Management. I have seen the rate of IT training going to zero. Think about specialized training, where if you make a mistake, the consequence of your errors, are talked about on the tv news of the evening.

The whole idea everybody is just planning to save their butt, using these strings coming out of these numeric matrices, while suspending judgement, just shudders me in horror. A bit like those South Asia Airline companies, that were forbidding their pilots from landing airplanes with manual piloting, leading to an increase loss of skills causing some well known disasters...

If well paid consultants cant even bother to check their links...

Re: EY Canada published a cybersecurity report and most citations were hallucinated

#116

I don't quite get it why they can't take another LLM and vet the output of the first with the second one. Surely they would not have the same hallucinations and would be able to detect hallucinations of the earlier LLM. Maybe it would cost too much in terms of tokens? I don't know but I would expect it to be realtively easy for an LLM to detect "hallucinations".

I am not exactly sure if this would solve the overall problem. The main one being lack of oversight. The solution to a social issue generally isn’t to throw more technology at it.

IBM once said “a computer can never be held accountable. Therefore a computer must never make a management decision”

Re: EY Canada published a cybersecurity report and most citations were hallucinated

#117
post #105

Earlier quoted context omitted.

The trope about external consultants is that your VP brings them in to review the company, and they talk to everybody and write a report on how to improve the business, and the report says exactly what you've been telling your VP but they've been ignoring you.

You are closer to the truth :) they are not simply paid to do nothing. They are paid to do dirty work.

They are paid to justify decisions executives have already made. It's often referred to as due diligence, but in practice these reports mostly just allow executives to tell the board it wasn't their fault if it goes wrong.

Re: EY Canada published a cybersecurity report and most citations were hallucinated

#119
post #8

The problem we're seeing across many professions is AI output is not getting vetted by knowledgeable people, whether it's an experienced analyst, senior engineer, expert attorney, or the resident physician. At best they skim, at worst they don't even see it at all before it's published, pushed to production, distributed to clients, or submitted to the court. In many cases the skills are available in house to do the n…

As an attorney, I feel like vetting AI output takes longer than just doing it from scratch, let alone versus just using a traditional form. With AI, I have to read through everything, often explain why it's wrong, and then rewrite everything anyways. I mean, I get way more billables, but I think it's symptomatic of how AI loses its advantage of being quick and accessible to those who don't understand the subject matt…

Be afraid, be very afraid:

"AI Hallucination Cases" - https://www.damiencharlotin.com/hallucinations/

Post reply on HN