If that is the case with a task so simple, why would we rely on these tools for high risk applications like medical diagnosis or analyzing financial data?
AI assistants misrepresent news content 45% of the time
31–40 of 306 posts
Re: AI assistants misrepresent news content 45% of the time
#32> This time, we used the free/consumer versions of ChatGPT, Copilot, Perplexity and Gemini.
IOW, they tested ChatGPT twice (Copilot uses ChatGPT's models) and didn't test Grok (or others).
Re: AI assistants misrepresent news content 45% of the time
#33Now let's run this experiment against the editorial boards in newsrooms. Obviously, AI isn't an improvement, but people who blindly trust the news have always been credulous rubes. It's just that the alternative is being completely ignorant of the worldviews of everyone around you. Peer-reviewed science is as close as we can get to good consensus and there's a lot of reasons this doesn't work for reporting.
Or against people in general.
It's a pet peeve of mine that we get these kinds of articles without a baseline established of how people do on the same measure.
Is misrepresenting news content 45% of the time better or worse than the average person? I don't know.
By extension: Would a person using an AI assistant misrepresent news more or less after having read a summary of the news provided by an AI assistant? I don't know that either.
When they have a "Why this distortion matters" section, those things matter. They've not established if this will make things better or worse.
(the cynic in me want another question answered too: How often does reporters misrepresent the news? Would it be better or worse if AI reviewed the facts and presented them vs. letting reporters do it? again: no idea)
Re: AI assistants misrepresent news content 45% of the time
#34According to PEW that's about the same % that trust the BBC's reporting. https://www.pewresearch.org/journalism/fact-sheet/news-media...
Re: AI assistants misrepresent news content 45% of the time
#35According to PEW that's about the same % that trust the BBC's reporting. https://www.pewresearch.org/journalism/fact-sheet/news-media...
Re: AI assistants misrepresent news content 45% of the time
#36It's important to bear this in mind whenever you find out that someone uses an LLM to summarize a meeting, email, or other communication you've held. That person is not really getting the message you were conveying.
That's a scary thought to me. They're not just outsourcing their thinking. They are actively sabotaging the only tool in their arsenal that could ever supplant it. I've felt it myself. Recently I was looking as some documentation without a clear edit history. I thought about feeding it into an AI and having it generate one for me, but didn't because I didn't have the time. To think, if I had done that, it probably wo…
Re: AI assistants misrepresent news content 45% of the time
#37> 45% of all AI answers had at least one significant issue. > 31% of responses showed serious sourcing problems – missing, misleading, or incorrect attributions. > 20% contained major accuracy issues, including hallucinated details and outdated information. I'm generally against whataboutism, but here I think we absolutely have to compare it to human-written news reports. Famously, Michael Crichton introduced the "Ge…
The problem highlighted here is that AI summaries misrepresent the original stories. This just opens a flood gate of slop that is 45% worse than the source, which wasn't stellar to begin with as you point out.
It's also not clear if humans do better when consuming either, and whether the effect of an AI summary, even with substantial issues, is to make the human reading them better or worse informed.
E.g. if it helps a person digest more material by getting more focused reports, it's entirely possible that flawed summaries would still in aggregate lead to a better understanding of a subject.
On its own, this article is just pure sensationalism.
Re: AI assistants misrepresent news content 45% of the time
#38Other issues: the report doesn't even say which particular models it's querying [ETA: discovered they do list this in an appendix], aside from saying it's the consumer tier. And it leaves off Anthropic (in my experience, by far the best at this type of task), favoring Perplexity and (perplexingly) Copilot. The article also intermingles claims from the recent report and the one on research conducted a year ago, leaving out critical context that... things have changed.
This article contains significant issues.
Re: AI assistants misrepresent news content 45% of the time
#39> ChatGPT / CBC / Is Türkiye in the EU?
> ChatGPT linked to a non-existent Wikipedia article on the “European Union Enlargement Goals for 2040”. In fact, there is no official EU policy under that name. The response hallucinates a URL but also, indirectly, an EU goal and policy.