Live data from Hacker News

AI assistants misrepresent news content 45% of the time

bbc.co.uk

51–60 of 306 posts

Re: AI assistants misrepresent news content 45% of the time

#51
post #42
post #38

If you dig into the actual report (I know, I know, how passe), you see how they get the numbers. Most of the errors are "sourcing issues": the AI assistant doesn't cite a claim, or it (shocking) cites Wikipedia instead of the BBC. Other issues: the report doesn't even say which particular models it's querying [ETA: discovered they do list this in an appendix], aside from saying it's the consumer tier. And it leaves o…

> or it (shocking) cites Wikipedia instead of the BBC. No... the problem is that it cites Wikipedia articles that don't exist . > ChatGPT linked to a non-existent Wikipedia article on the “European Union Enlargement Goals for 2040”. In fact, there is no official EU policy under that name. The response hallucinates a URL but also, indirectly, an EU goal and policy.

Actually there was a Wikipedia article of this name, but it was deleted in June -- because it was AI generated. Unfortunately AI falls for this much like humans do.

https://en.wikipedia.org/wiki/Wikipedia:Articles_for_deletio...

Re: AI assistants misrepresent news content 45% of the time

#52
post #42
post #38

If you dig into the actual report (I know, I know, how passe), you see how they get the numbers. Most of the errors are "sourcing issues": the AI assistant doesn't cite a claim, or it (shocking) cites Wikipedia instead of the BBC. Other issues: the report doesn't even say which particular models it's querying [ETA: discovered they do list this in an appendix], aside from saying it's the consumer tier. And it leaves o…

> or it (shocking) cites Wikipedia instead of the BBC. No... the problem is that it cites Wikipedia articles that don't exist . > ChatGPT linked to a non-existent Wikipedia article on the “European Union Enlargement Goals for 2040”. In fact, there is no official EU policy under that name. The response hallucinates a URL but also, indirectly, an EU goal and policy.

> For the current research, a set of 30 “core” news questions was developed

Right. Let's talk about statistics for a bit. Or let's put it differently: they found in their report that 45% of the answers for 30 questions they have "developed" had a significant issue, e.g. inexisting reference

I'll give you 30 questions out of my sleeve where 95% of the answers will not have any significant issue.

Re: AI assistants misrepresent news content 45% of the time

#53
post #15

Earlier quoted context omitted.

Human news isn't a good comparison because this is second order -- LMMs are downstream of human news. It's a game of stochastic telephone. All the human error is carried through with additional hallucinations on top.

But the issue is that the vast majority of "human news" is second order (at best), essentially paraphrasing releases by news agencies like Reuters or Associated Press, or scientific articles, and typically doing a horrible job at it. Regarding scientific reporting, there's as usual a relevant xkcd ("New Study") [0], and in this case even better, there's a fabulous one from PhD Comics ("Science News Cycle") [1]. [0] h…

Then the point still stands, this makes things even worse given that it's adding its own hallucinations on top, instead of simply relaying the content or idealistically, identifying issues in the reporting.

Re: AI assistants misrepresent news content 45% of the time

#54
post #15

Earlier quoted context omitted.

Human news isn't a good comparison because this is second order -- LMMs are downstream of human news. It's a game of stochastic telephone. All the human error is carried through with additional hallucinations on top.

But the issue is that the vast majority of "human news" is second order (at best), essentially paraphrasing releases by news agencies like Reuters or Associated Press, or scientific articles, and typically doing a horrible job at it. Regarding scientific reporting, there's as usual a relevant xkcd ("New Study") [0], and in this case even better, there's a fabulous one from PhD Comics ("Science News Cycle") [1]. [0] h…

You understand that an LLM can only poorly regurgitate whatever it’s fed right? An LLM will _always_ be less useful than a primary/secondary source, because they can’t fucking think.

Re: AI assistants misrepresent news content 45% of the time

#56
post #12

Kagi News has been pretty accurate. Source information is provided along with the summary and key details too. AI summarizes are good for getting a feel of if you want to read an article or not. Even with Kagi News I verify key facts myself.

What if the AI makes an interesting or important article sound like one you don't want to read? You'd never cross check the fact, and you'd never discover how wrong the AI was.

Integrity of words and author intent is important. I understand the intent of your hypothetical but I haven’t run into this issue in practice with Kagi News.

Never share information about an article you have not read. Likewise, never draw definitive conclusions from an article that is not of interest.

If you do not find a headline interesting, the take away is that you did not find the headline interesting. Nothing more, nothing less. You should read the key insights before dismissing an article entirely.

I can imagine AI summarizes being problematic for a class of people that do not cross check if an article is of value to them.

Re: AI assistants misrepresent news content 45% of the time

#57
post #42
post #38

If you dig into the actual report (I know, I know, how passe), you see how they get the numbers. Most of the errors are "sourcing issues": the AI assistant doesn't cite a claim, or it (shocking) cites Wikipedia instead of the BBC. Other issues: the report doesn't even say which particular models it's querying [ETA: discovered they do list this in an appendix], aside from saying it's the consumer tier. And it leaves o…

> or it (shocking) cites Wikipedia instead of the BBC. No... the problem is that it cites Wikipedia articles that don't exist . > ChatGPT linked to a non-existent Wikipedia article on the “European Union Enlargement Goals for 2040”. In fact, there is no official EU policy under that name. The response hallucinates a URL but also, indirectly, an EU goal and policy.

Do we have any good research on how much less often larger, newer models will just make stuff up like this? As it is, it's pretty clear LLMs are categorically not a good idea for directly querying for information in any non-fiction-writing context. If you're using an LLM to research something that needs to be accurate, the LLM needs to be doing a tool call to a web search and only asked to summarize relevant facts from the existing information it can find, and have them be cited by hard-coding the UI to link the pages the LLM reviewed. The LLM itself cannot be trusted to generate its own citations. It will just generate something that looks like a relevant citation, along with whatever imaginary content it wants to attribute to this non-existent source.

Re: AI assistants misrepresent news content 45% of the time

#58
post #9

> 45% of all AI answers had at least one significant issue. > 31% of responses showed serious sourcing problems – missing, misleading, or incorrect attributions. > 20% contained major accuracy issues, including hallucinated details and outdated information. I'm generally against whataboutism, but here I think we absolutely have to compare it to human-written news reports. Famously, Michael Crichton introduced the "Ge…

That's not comparable. Reading news reports and summarizing them is about a thousand times easier than writing those news reports in the first place. If you want to see how humans fare at this task, have some people answer questions about the news and then compare their answers to the original reporting. I'm not sure if the average human would fare too well at this either, but it's completely different from the question of how accurate the original news itself is.

Re: AI assistants misrepresent news content 45% of the time

#59
The media today is so polarized, so dishonest, and so bent on feeding the egos of it's users, the bar to pass them is literally underground.

You can go through most big name media stories and find it ridden with omissions of uncomfortable facts, careful structuring of words to give the illusion of untrue facts being true, and careful curation of what stories are reported.

More than anything, I hope AI topples the garbage bin fire that is modern "journalism". Also, it should be very clear why the media is especially hostile towards AI. It might reveal them as the clowns they are, and kill the social division and controversy that is their lifeblood.

Re: AI assistants misrepresent news content 45% of the time

#60
post #20

I recently tried to get Gemini to collect fresh news and show them to me, and instead of using search it hallucinated everything wholesale, titles, abstracts and links. Not just once, multiple times. I am kind of afraid of using Gemini now for anything related to web search. Here is a sample: > [1] Google DeepMind and Harvard researchers propose a new method for testing the ‘theory of mind’ of LLMs - Researchers have…

They can be good for search, but you must click through the provided links and verify that they actually say what it says they do.

The problem is that 90% of people will not do that once they've satisfied their confirmation bias. Hard to say if that's going to be better or worse than the current echo chamber effects of the Internet. I'm still holding out for better, but certainly this is shaking that assumption
Post reply on HN