"AI assistants misrepresent news content 45% of the time" How does that compare to the number for reporters? I feel like half the time I read or hear a report on a subject I know the reporter misrepresented something.
That’s whataboutism and doesn’t address the criticism or the problem. If a reporter misrepresents a subject, intentionally or accidentally, it doesn’t make it OK for a tool to then misrepresent it further, mangling both was correct and what was incorrect. https://en.wikipedia.org/wiki/Whataboutism
AI assistants misrepresent news content 45% of the time
221–230 of 306 posts
Re: AI assistants misrepresent news content 45% of the time
#222Earlier quoted context omitted.
It's probably for the best that chat interfaces avoid making direct HTTP calls to sources at run-time to confirm that they don't 404 - imagine how much extra traffic that could add to an internet ecosystem which is suffering from badly written crawlers already. (Not to mention plenty of sites have added robots.txt rules deliberately excluding known AI user-agents now.)
Wouldn't it be the same amount of requests as a regular person researching something the old way?
Re: AI assistants misrepresent news content 45% of the time
#223Re: AI assistants misrepresent news content 45% of the time
#224> All participating organizations then generated responses to each question from each of the four AI assistants. This time, we used the free/consumer versions of ChatGPT, Copilot, Perplexity and Gemini. Free versions were chosen to replicate the default (and likely most common) experience for users. Responses were generated in late May and early June 2025. First of all, none of the SOTA models we're currently using w…
Ah, the "you're using the wrong model" fallacy (is there a name for this?) In the eyes of the evangelists, every major model seems to go from " This model is close to flawless at this task, you MUST try this TODAY " to " It's absolutely wild that anyone would ever consider using such a no-good, worthless model for this task " over the course of a year or so. The old model has to be re-framed for the new model to look…
What I’m saying is that there should be a disclaimer: hey, we’re testing these models for the average person, that have no idea about AI. People who actually know AI would never use them in this way.
A better idea: educate people. Add “Here’s the best way to use them btw…” to the report.
All I’m saying is, it’s a tool, and yes you can use it wrong. That’s not a crazy realization. It applies to every other tool.
We knew that the hallucation rate for gpt 4o was nuts. From the start. We also know that gpt-5 has a much lower hallucination rate. So there are no surprises here, I’m not saying anything groundbreaking, and neither are they.
Re: AI assistants misrepresent news content 45% of the time
#225Earlier quoted context omitted.
That’s whataboutism and doesn’t address the criticism or the problem. If a reporter misrepresents a subject, intentionally or accidentally, it doesn’t make it OK for a tool to then misrepresent it further, mangling both was correct and what was incorrect. https://en.wikipedia.org/wiki/Whataboutism
It is not OK, but if it's lower, it is an improvement.
Re: AI assistants misrepresent news content 45% of the time
#226If you dig into the actual report (I know, I know, how passe), you see how they get the numbers. Most of the errors are "sourcing issues": the AI assistant doesn't cite a claim, or it (shocking) cites Wikipedia instead of the BBC. Other issues: the report doesn't even say which particular models it's querying [ETA: discovered they do list this in an appendix], aside from saying it's the consumer tier. And it leaves o…
> or it (shocking) cites Wikipedia instead of the BBC. No... the problem is that it cites Wikipedia articles that don't exist . > ChatGPT linked to a non-existent Wikipedia article on the “European Union Enlargement Goals for 2040”. In fact, there is no official EU policy under that name. The response hallucinates a URL but also, indirectly, an EU goal and policy.
Re: AI assistants misrepresent news content 45% of the time
#227> All participating organizations then generated responses to each question from each of the four AI assistants. This time, we used the free/consumer versions of ChatGPT, Copilot, Perplexity and Gemini. Free versions were chosen to replicate the default (and likely most common) experience for users. Responses were generated in late May and early June 2025. First of all, none of the SOTA models we're currently using w…
If they used a paid version, their study would not represent how most people use AI (with the free version)
Re: AI assistants misrepresent news content 45% of the time
#228> All participating organizations then generated responses to each question from each of the four AI assistants. This time, we used the free/consumer versions of ChatGPT, Copilot, Perplexity and Gemini. Free versions were chosen to replicate the default (and likely most common) experience for users. Responses were generated in late May and early June 2025. First of all, none of the SOTA models we're currently using w…
I think you are missing the point: it's mainly to highlight that the models that most people use, i.e. free versions with default settings, output a large number of factual errors, even when they are asked to base their answer to specific sources of information (as it's explained in their methodology document).
Re: AI assistants misrepresent news content 45% of the time
#229> All participating organizations then generated responses to each question from each of the four AI assistants. This time, we used the free/consumer versions of ChatGPT, Copilot, Perplexity and Gemini. Free versions were chosen to replicate the default (and likely most common) experience for users. Responses were generated in late May and early June 2025. First of all, none of the SOTA models we're currently using w…
> On top of that, to use free models for anything like this is absolutely wild. I use the absolute best models, and the research versions of this whenever I do research. Anything less is inviting disaster. "I contend we are both atheists, I just believe in one fewer god than you do. When you understand why you dismiss all the other possible gods, you will understand why I dismiss yours." - Stephen F Roberts
One knife does not cut potatoes. Doesn’t mean that all knives don’t cut potatoes. Use the right tool for the job.
Though I do love a well placed quote
Re: AI assistants misrepresent news content 45% of the time
#230Earlier quoted context omitted.
> It's a pet peeve of mine that we get these kinds of articles without a baseline established of how people do on the same measure I don’t have a personal human news summarizer? The comparison is between a human reading the primary source against the same human reading an LLM hallucination mixed with an LLM referring the primary source. > cynic in me want another question answered too: How often does reporters misrep…
> I don’t have a personal human news summarizer? Not a personal one. You do however have reporters sitting between you and the source material a lot of the time, and sometimes multiple levels of reporters playing games of telephone with the source material. > The comparison is between a human reading the primary source against the same human reading an LLM hallucination mixed with an LLM referring the primary source.…
In cases where a reporter is just summarising e.g. a court case, sure. Stock market news has been automated since the 2000s.
More broadly, AI assistants misrepresenting news content may sometimes direct reference a court case. But they often don't. Even if they only could, that covers a small fraction of the news, much of which the AI will need to rely on reporters detailing the primary sources they're interfacing with.
Reporter error is somewhat orthogonal to AI assistants' accuracy.