Live data from Hacker News

AI assistants misrepresent news content 45% of the time

bbc.co.uk

221–230 of 306 posts

Re: AI assistants misrepresent news content 45% of the time

#221
post #180

"AI assistants misrepresent news content 45% of the time" How does that compare to the number for reporters? I feel like half the time I read or hear a report on a subject I know the reporter misrepresented something.

That’s whataboutism and doesn’t address the criticism or the problem. If a reporter misrepresents a subject, intentionally or accidentally, it doesn’t make it OK for a tool to then misrepresent it further, mangling both was correct and what was incorrect. https://en.wikipedia.org/wiki/Whataboutism

It is not OK, but if it's lower, it is an improvement.

Re: AI assistants misrepresent news content 45% of the time

#222
post #146

Earlier quoted context omitted.

It's probably for the best that chat interfaces avoid making direct HTTP calls to sources at run-time to confirm that they don't 404 - imagine how much extra traffic that could add to an internet ecosystem which is suffering from badly written crawlers already. (Not to mention plenty of sites have added robots.txt rules deliberately excluding known AI user-agents now.)

Wouldn't it be the same amount of requests as a regular person researching something the old way?

If you watch the thinking panel in ChatGPT with GPT-5 Thinking it often consults dozens of pages in response to a single prompt.

Re: AI assistants misrepresent news content 45% of the time

#224
post #63

> All participating organizations then generated responses to each question from each of the four AI assistants. This time, we used the free/consumer versions of ChatGPT, Copilot, Perplexity and Gemini. Free versions were chosen to replicate the default (and likely most common) experience for users. Responses were generated in late May and early June 2025. First of all, none of the SOTA models we're currently using w…

Ah, the "you're using the wrong model" fallacy (is there a name for this?) In the eyes of the evangelists, every major model seems to go from " This model is close to flawless at this task, you MUST try this TODAY " to " It's absolutely wild that anyone would ever consider using such a no-good, worthless model for this task " over the course of a year or so. The old model has to be re-framed for the new model to look…

Not an evangelist for AI at all, I just love it as a tool for my creativity, research and coding.

What I’m saying is that there should be a disclaimer: hey, we’re testing these models for the average person, that have no idea about AI. People who actually know AI would never use them in this way.

A better idea: educate people. Add “Here’s the best way to use them btw…” to the report.

All I’m saying is, it’s a tool, and yes you can use it wrong. That’s not a crazy realization. It applies to every other tool.

We knew that the hallucation rate for gpt 4o was nuts. From the start. We also know that gpt-5 has a much lower hallucination rate. So there are no surprises here, I’m not saying anything groundbreaking, and neither are they.

Re: AI assistants misrepresent news content 45% of the time

#225
post #180

Earlier quoted context omitted.

That’s whataboutism and doesn’t address the criticism or the problem. If a reporter misrepresents a subject, intentionally or accidentally, it doesn’t make it OK for a tool to then misrepresent it further, mangling both was correct and what was incorrect. https://en.wikipedia.org/wiki/Whataboutism

It is not OK, but if it's lower, it is an improvement.

It can’t be lower. LLMs work on the text they’re given. The submission isn’t saying that LLMs misrepresent half of reality, but of the news content they consume. In other words, even if news sources have errors, LLMs are adding to them.

Re: AI assistants misrepresent news content 45% of the time

#226
post #42
post #38

If you dig into the actual report (I know, I know, how passe), you see how they get the numbers. Most of the errors are "sourcing issues": the AI assistant doesn't cite a claim, or it (shocking) cites Wikipedia instead of the BBC. Other issues: the report doesn't even say which particular models it's querying [ETA: discovered they do list this in an appendix], aside from saying it's the consumer tier. And it leaves o…

> or it (shocking) cites Wikipedia instead of the BBC. No... the problem is that it cites Wikipedia articles that don't exist . > ChatGPT linked to a non-existent Wikipedia article on the “European Union Enlargement Goals for 2040”. In fact, there is no official EU policy under that name. The response hallucinates a URL but also, indirectly, an EU goal and policy.

Existing is just a point in time

Re: AI assistants misrepresent news content 45% of the time

#227
post #63

> All participating organizations then generated responses to each question from each of the four AI assistants. This time, we used the free/consumer versions of ChatGPT, Copilot, Perplexity and Gemini. Free versions were chosen to replicate the default (and likely most common) experience for users. Responses were generated in late May and early June 2025. First of all, none of the SOTA models we're currently using w…

If they used a paid version, their study would not represent how most people use AI (with the free version)

But they’re using a free version that’s not even out there anymore. This is my problem - it came out already dated.

Re: AI assistants misrepresent news content 45% of the time

#228
post #69
post #63

> All participating organizations then generated responses to each question from each of the four AI assistants. This time, we used the free/consumer versions of ChatGPT, Copilot, Perplexity and Gemini. Free versions were chosen to replicate the default (and likely most common) experience for users. Responses were generated in late May and early June 2025. First of all, none of the SOTA models we're currently using w…

I think you are missing the point: it's mainly to highlight that the models that most people use, i.e. free versions with default settings, output a large number of factual errors, even when they are asked to base their answer to specific sources of information (as it's explained in their methodology document).

Is it true of the latest free models? Just saying that the report started already dated.

Re: AI assistants misrepresent news content 45% of the time

#229
post #63

> All participating organizations then generated responses to each question from each of the four AI assistants. This time, we used the free/consumer versions of ChatGPT, Copilot, Perplexity and Gemini. Free versions were chosen to replicate the default (and likely most common) experience for users. Responses were generated in late May and early June 2025. First of all, none of the SOTA models we're currently using w…

> On top of that, to use free models for anything like this is absolutely wild. I use the absolute best models, and the research versions of this whenever I do research. Anything less is inviting disaster. "I contend we are both atheists, I just believe in one fewer god than you do. When you understand why you dismiss all the other possible gods, you will understand why I dismiss yours." - Stephen F Roberts

It ain’t a God, it’s a tool.

One knife does not cut potatoes. Doesn’t mean that all knives don’t cut potatoes. Use the right tool for the job.

Though I do love a well placed quote

Re: AI assistants misrepresent news content 45% of the time

#230
post #160

Earlier quoted context omitted.

> It's a pet peeve of mine that we get these kinds of articles without a baseline established of how people do on the same measure I don’t have a personal human news summarizer? The comparison is between a human reading the primary source against the same human reading an LLM hallucination mixed with an LLM referring the primary source. > cynic in me want another question answered too: How often does reporters misrep…

> I don’t have a personal human news summarizer? Not a personal one. You do however have reporters sitting between you and the source material a lot of the time, and sometimes multiple levels of reporters playing games of telephone with the source material. > The comparison is between a human reading the primary source against the same human reading an LLM hallucination mixed with an LLM referring the primary source.…

> You do however have reporters sitting between you and the source material a lot of the time

In cases where a reporter is just summarising e.g. a court case, sure. Stock market news has been automated since the 2000s.

More broadly, AI assistants misrepresenting news content may sometimes direct reference a court case. But they often don't. Even if they only could, that covers a small fraction of the news, much of which the AI will need to rely on reporters detailing the primary sources they're interfacing with.

Reporter error is somewhat orthogonal to AI assistants' accuracy.

Post reply on HN