Live data from Hacker News

AI assistants misrepresent news content 45% of the time

bbc.co.uk

31–40 of 306 posts

Re: AI assistants misrepresent news content 45% of the time

#31
One thing that makes me pessimistic about the short term utility of LLMs has been their inability to produce basic media monitoring documents. This is an intern type entry level task that it simply cannot complete with any reliability or consistency. It doesn't matter if I use the expensive paid services or spend dozens of prompts trying to configure, it simply wont produce a document that is of any use to me.

If that is the case with a task so simple, why would we rely on these tools for high risk applications like medical diagnosis or analyzing financial data?

Re: AI assistants misrepresent news content 45% of the time

#33

Now let's run this experiment against the editorial boards in newsrooms. Obviously, AI isn't an improvement, but people who blindly trust the news have always been credulous rubes. It's just that the alternative is being completely ignorant of the worldviews of everyone around you. Peer-reviewed science is as close as we can get to good consensus and there's a lot of reasons this doesn't work for reporting.

> Now let's run this experiment against the editorial boards in newsrooms.

Or against people in general.

It's a pet peeve of mine that we get these kinds of articles without a baseline established of how people do on the same measure.

Is misrepresenting news content 45% of the time better or worse than the average person? I don't know.

By extension: Would a person using an AI assistant misrepresent news more or less after having read a summary of the news provided by an AI assistant? I don't know that either.

When they have a "Why this distortion matters" section, those things matter. They've not established if this will make things better or worse.

(the cynic in me want another question answered too: How often does reporters misrepresent the news? Would it be better or worse if AI reviewed the facts and presented them vs. letting reporters do it? again: no idea)

Re: AI assistants misrepresent news content 45% of the time

#36

It's important to bear this in mind whenever you find out that someone uses an LLM to summarize a meeting, email, or other communication you've held. That person is not really getting the message you were conveying.

That's a scary thought to me. They're not just outsourcing their thinking. They are actively sabotaging the only tool in their arsenal that could ever supplant it. I've felt it myself. Recently I was looking as some documentation without a clear edit history. I thought about feeding it into an AI and having it generate one for me, but didn't because I didn't have the time. To think, if I had done that, it probably wo…

You've gotta be careful using "not just X, but Y" these days ;).

Re: AI assistants misrepresent news content 45% of the time

#37
post #18
post #9

> 45% of all AI answers had at least one significant issue. > 31% of responses showed serious sourcing problems – missing, misleading, or incorrect attributions. > 20% contained major accuracy issues, including hallucinated details and outdated information. I'm generally against whataboutism, but here I think we absolutely have to compare it to human-written news reports. Famously, Michael Crichton introduced the "Ge…

The problem highlighted here is that AI summaries misrepresent the original stories. This just opens a flood gate of slop that is 45% worse than the source, which wasn't stellar to begin with as you point out.

A whole lot of news is regurgiated wire service reports, so how reporters do matters greatly - if they're doing badly, then it's entirely possible that an AI summary of the wire service releases would be an improvement (probably not, but without a baseline we don't know)

It's also not clear if humans do better when consuming either, and whether the effect of an AI summary, even with substantial issues, is to make the human reading them better or worse informed.

E.g. if it helps a person digest more material by getting more focused reports, it's entirely possible that flawed summaries would still in aggregate lead to a better understanding of a subject.

On its own, this article is just pure sensationalism.

Re: AI assistants misrepresent news content 45% of the time

#38
If you dig into the actual report (I know, I know, how passe), you see how they get the numbers. Most of the errors are "sourcing issues": the AI assistant doesn't cite a claim, or it (shocking) cites Wikipedia instead of the BBC.

Other issues: the report doesn't even say which particular models it's querying [ETA: discovered they do list this in an appendix], aside from saying it's the consumer tier. And it leaves off Anthropic (in my experience, by far the best at this type of task), favoring Perplexity and (perplexingly) Copilot. The article also intermingles claims from the recent report and the one on research conducted a year ago, leaving out critical context that... things have changed.

This article contains significant issues.

Re: AI assistants misrepresent news content 45% of the time

#39
Page 10 onwards of this PDF shows concrete examples of the mistakes: https://www.bbc.co.uk/aboutthebbc/documents/news-integrity-i...

> ChatGPT / CBC / Is Türkiye in the EU?

> ChatGPT linked to a non-existent Wikipedia article on the “European Union Enlargement Goals for 2040”. In fact, there is no official EU policy under that name. The response hallucinates a URL but also, indirectly, an EU goal and policy.

Post reply on HN