Live data from Hacker News

AI assistants misrepresent news content 45% of the time

bbc.co.uk

61–70 of 306 posts

Re: AI assistants misrepresent news content 45% of the time

#62
post #42

Earlier quoted context omitted.

> or it (shocking) cites Wikipedia instead of the BBC. No... the problem is that it cites Wikipedia articles that don't exist . > ChatGPT linked to a non-existent Wikipedia article on the “European Union Enlargement Goals for 2040”. In fact, there is no official EU policy under that name. The response hallucinates a URL but also, indirectly, an EU goal and policy.

> For the current research, a set of 30 “core” news questions was developed Right. Let's talk about statistics for a bit. Or let's put it differently: they found in their report that 45% of the answers for 30 questions they have "developed" had a significant issue, e.g. inexisting reference I'll give you 30 questions out of my sleeve where 95% of the answers will not have any significant issue.

Yes, I'm sure you could hack together some bullshit questions to demonstrate whatever you want. Is there a specific reason that the reasonably straightforward methodology they did use is somehow flawed?

Re: AI assistants misrepresent news content 45% of the time

#63
> All participating organizations then generated responses to each question from each of the four AI assistants. This time, we used the free/consumer versions of ChatGPT, Copilot, Perplexity and Gemini. Free versions were chosen to replicate the default (and likely most common) experience for users. Responses were generated in late May and early June 2025.

First of all, none of the SOTA models we're currently using were released in May and early June. Gemini 2.5 came out in June 17, GPT 5 & Claude Opus 4.1 at the beginning of August.

On top of that, to use free models for anything like this is absolutely wild. I use the absolute best models, and the research versions of this whenever I do research. Anything less is inviting disaster.

You have to use the right tools for the right job, and any report that is more than a month old is useless in the AI world at this point in time, beyond a snapshot of how things 'used to be'.

Re: AI assistants misrepresent news content 45% of the time

#64
post #12

Kagi News has been pretty accurate. Source information is provided along with the summary and key details too. AI summarizes are good for getting a feel of if you want to read an article or not. Even with Kagi News I verify key facts myself.

What if the AI makes an interesting or important article sound like one you don't want to read? You'd never cross check the fact, and you'd never discover how wrong the AI was.

That's fair, but i also don't cross check news sources on average either. I should, but there in lies the real problem imo. Information is war these days, and we've not yet developed tools for wading through immense piles of subtly inaccurate or biased data.

We're in a weird time. It's always been like this, it's just much.. more, now. I'm not sure how we'll adapt.

Re: AI assistants misrepresent news content 45% of the time

#66
post #42

Earlier quoted context omitted.

> or it (shocking) cites Wikipedia instead of the BBC. No... the problem is that it cites Wikipedia articles that don't exist . > ChatGPT linked to a non-existent Wikipedia article on the “European Union Enlargement Goals for 2040”. In fact, there is no official EU policy under that name. The response hallucinates a URL but also, indirectly, an EU goal and policy.

Actually there was a Wikipedia article of this name, but it was deleted in June -- because it was AI generated. Unfortunately AI falls for this much like humans do. https://en.wikipedia.org/wiki/Wikipedia:Articles_for_deletio...

This is likely because of the knowledge cutoff.

I have seen a few cases before of "hallucinations" that turned out to be things that did exist, but no longer do.

Re: AI assistants misrepresent news content 45% of the time

#68

I am reading the actual report and some of this seems _quite_ nitpicky: > ChatGPT / Radio-Canada / Is Trump starting a trade war? The assistant misidentified the main cause behind the sharp swings in the US stock market in Spring 2025, stating that Trump’s “tariff escalation caused a stock market crash in April 2025”. As RadioCanada’s evaluator notes: “In fact it was not the escalation between Washington and its Nort…

That may be nitpicky, but I don't think it's too much to ask that a computer system be fully factually accurate when it comes to basic objective numerical facts. This is very much a case of, "if it gets this stuff wrong, what else is it getting wrong?"

Re: AI assistants misrepresent news content 45% of the time

#69
post #63

> All participating organizations then generated responses to each question from each of the four AI assistants. This time, we used the free/consumer versions of ChatGPT, Copilot, Perplexity and Gemini. Free versions were chosen to replicate the default (and likely most common) experience for users. Responses were generated in late May and early June 2025. First of all, none of the SOTA models we're currently using w…

I think you are missing the point: it's mainly to highlight that the models that most people use, i.e. free versions with default settings, output a large number of factual errors, even when they are asked to base their answer to specific sources of information (as it's explained in their methodology document).
Post reply on HN