Live data from Hacker News

AI assistants misrepresent news content 45% of the time

bbc.co.uk

121–130 of 306 posts

Re: AI assistants misrepresent news content 45% of the time

#121
post #42
post #38

If you dig into the actual report (I know, I know, how passe), you see how they get the numbers. Most of the errors are "sourcing issues": the AI assistant doesn't cite a claim, or it (shocking) cites Wikipedia instead of the BBC. Other issues: the report doesn't even say which particular models it's querying [ETA: discovered they do list this in an appendix], aside from saying it's the consumer tier. And it leaves o…

> or it (shocking) cites Wikipedia instead of the BBC. No... the problem is that it cites Wikipedia articles that don't exist . > ChatGPT linked to a non-existent Wikipedia article on the “European Union Enlargement Goals for 2040”. In fact, there is no official EU policy under that name. The response hallucinates a URL but also, indirectly, an EU goal and policy.

I used perplexity for searches and I clicked on all sources that were given. Depending on the model used from 100% to 20% of the urls I tested did not exist. I kept on querying the LLM about it and it finally told me that it generated « the most probable » urls for the topic in question based on the ones he knows exists. Useless.

Re: AI assistants misrepresent news content 45% of the time

#122
post #42

Earlier quoted context omitted.

> or it (shocking) cites Wikipedia instead of the BBC. No... the problem is that it cites Wikipedia articles that don't exist . > ChatGPT linked to a non-existent Wikipedia article on the “European Union Enlargement Goals for 2040”. In fact, there is no official EU policy under that name. The response hallucinates a URL but also, indirectly, an EU goal and policy.

Do we have any good research on how much less often larger, newer models will just make stuff up like this? As it is, it's pretty clear LLMs are categorically not a good idea for directly querying for information in any non-fiction-writing context. If you're using an LLM to research something that needs to be accurate, the LLM needs to be doing a tool call to a web search and only asked to summarize relevant facts fr…

The problem is that people are using it as a substitute for a web search, and the web search company has decided to kill off search as a product and pivot to video, err, I mean pivot to AI chatbots so hard they replaced one of the common ways to access emergency services on their mobile phones with an AI chatbot that can't help you in an emergency.

Not to mention, the AI companies have been extremely abusive to the rest of the internet so they are often blocked from accessing various web sites, so it's not like they're going to be able to access legitimate information anyways.

Re: AI assistants misrepresent news content 45% of the time

#123
post #63

> All participating organizations then generated responses to each question from each of the four AI assistants. This time, we used the free/consumer versions of ChatGPT, Copilot, Perplexity and Gemini. Free versions were chosen to replicate the default (and likely most common) experience for users. Responses were generated in late May and early June 2025. First of all, none of the SOTA models we're currently using w…

> On top of that, to use free models for anything like this is absolutely wild. I use the absolute best models, and the research versions of this whenever I do research. Anything less is inviting disaster.

"I contend we are both atheists, I just believe in one fewer god than you do. When you understand why you dismiss all the other possible gods, you will understand why I dismiss yours." - Stephen F Roberts

Re: AI assistants misrepresent news content 45% of the time

#124

Earlier quoted context omitted.

Yes, and you answered it yourself.

Err, no? Being _possible_ does not necessarily imply that's what happened.

A bucket of 30 questions is not a statistically significant sample size which we can use to support the hypothesis which goes to say that all AI assistants they tested are 45% of the time wrong. That's not how science works.

Neither is my bucket of 30 questions statistcally significant but it goes to say that I can disprove their hypothesis just by giving them my sample.

I think that the report is being disingenious and I don't understand for what reasons. it's funny that they say "misrepresent" when that's exactly what they are doing.

Re: AI assistants misrepresent news content 45% of the time

#125

Earlier quoted context omitted.

This is likely because of the knowledge cutoff. I have seen a few cases before of "hallucinations" that turned out to be things that did exist, but no longer do.

The fix for this is for the AI to double-check all links before providing them to the user. I frequently ask ChatGPT to double check that references actually exist when it gives me them. It should be built in!

I thought people here hated it when LLMs made http requests?

Re: AI assistants misrepresent news content 45% of the time

#126
post #38

If you dig into the actual report (I know, I know, how passe), you see how they get the numbers. Most of the errors are "sourcing issues": the AI assistant doesn't cite a claim, or it (shocking) cites Wikipedia instead of the BBC. Other issues: the report doesn't even say which particular models it's querying [ETA: discovered they do list this in an appendix], aside from saying it's the consumer tier. And it leaves o…

I wouldn't even say BBC is a good source to cite. For foreign news, BBC is outright biased. Though I don't have any good suggestions for what an LLM should cite instead.

Re: AI assistants misrepresent news content 45% of the time

#127
post #38

If you dig into the actual report (I know, I know, how passe), you see how they get the numbers. Most of the errors are "sourcing issues": the AI assistant doesn't cite a claim, or it (shocking) cites Wikipedia instead of the BBC. Other issues: the report doesn't even say which particular models it's querying [ETA: discovered they do list this in an appendix], aside from saying it's the consumer tier. And it leaves o…

I wouldn't even say BBC is a good source to cite. For foreign news, BBC is outright biased. Though I don't have any good suggestions for what an LLM should cite instead.

Ground News or something similar that at least attempts to aggregate, compare ownership, bias, and factuality.

Imo at least

Re: AI assistants misrepresent news content 45% of the time

#128
post #42

Earlier quoted context omitted.

> or it (shocking) cites Wikipedia instead of the BBC. No... the problem is that it cites Wikipedia articles that don't exist . > ChatGPT linked to a non-existent Wikipedia article on the “European Union Enlargement Goals for 2040”. In fact, there is no official EU policy under that name. The response hallucinates a URL but also, indirectly, an EU goal and policy.

I used perplexity for searches and I clicked on all sources that were given. Depending on the model used from 100% to 20% of the urls I tested did not exist. I kept on querying the LLM about it and it finally told me that it generated « the most probable » urls for the topic in question based on the ones he knows exists. Useless.

I share your opinion on the results, but why would you trust the LLM explanation for why it does what it does?

Re: AI assistants misrepresent news content 45% of the time

#129
post #38

If you dig into the actual report (I know, I know, how passe), you see how they get the numbers. Most of the errors are "sourcing issues": the AI assistant doesn't cite a claim, or it (shocking) cites Wikipedia instead of the BBC. Other issues: the report doesn't even say which particular models it's querying [ETA: discovered they do list this in an appendix], aside from saying it's the consumer tier. And it leaves o…

I wouldn't even say BBC is a good source to cite. For foreign news, BBC is outright biased. Though I don't have any good suggestions for what an LLM should cite instead.

Reuters or AP IMO. Both take NPOV and accuracy very seriously. Reuters famously wouldn't even refer to the 9/11 hijackers as terrorists, as they wanted to remain as value-neutral as possible.

Re: AI assistants misrepresent news content 45% of the time

#130
post #38

If you dig into the actual report (I know, I know, how passe), you see how they get the numbers. Most of the errors are "sourcing issues": the AI assistant doesn't cite a claim, or it (shocking) cites Wikipedia instead of the BBC. Other issues: the report doesn't even say which particular models it's querying [ETA: discovered they do list this in an appendix], aside from saying it's the consumer tier. And it leaves o…

I wouldn't even say BBC is a good source to cite. For foreign news, BBC is outright biased. Though I don't have any good suggestions for what an LLM should cite instead.

The BBC has a strong right wing bias within the UK too.

There’s no such thing as unbiased.

Post reply on HN