Live data from Hacker News

AI assistants misrepresent news content 45% of the time

bbc.co.uk

11–20 of 306 posts

Re: AI assistants misrepresent news content 45% of the time

#14
post #9

> 45% of all AI answers had at least one significant issue. > 31% of responses showed serious sourcing problems – missing, misleading, or incorrect attributions. > 20% contained major accuracy issues, including hallucinated details and outdated information. I'm generally against whataboutism, but here I think we absolutely have to compare it to human-written news reports. Famously, Michael Crichton introduced the "Ge…

[deleted]

Re: AI assistants misrepresent news content 45% of the time

#15
post #9

> 45% of all AI answers had at least one significant issue. > 31% of responses showed serious sourcing problems – missing, misleading, or incorrect attributions. > 20% contained major accuracy issues, including hallucinated details and outdated information. I'm generally against whataboutism, but here I think we absolutely have to compare it to human-written news reports. Famously, Michael Crichton introduced the "Ge…

Human news isn't a good comparison because this is second order -- LMMs are downstream of human news. It's a game of stochastic telephone. All the human error is carried through with additional hallucinations on top.

Re: AI assistants misrepresent news content 45% of the time

#16
post #9

> 45% of all AI answers had at least one significant issue. > 31% of responses showed serious sourcing problems – missing, misleading, or incorrect attributions. > 20% contained major accuracy issues, including hallucinated details and outdated information. I'm generally against whataboutism, but here I think we absolutely have to compare it to human-written news reports. Famously, Michael Crichton introduced the "Ge…

I'd say the 45% is on top of mistakes by Journalists themselves. "AI" takes certain newspapers as gospel, and it is easy to find omissions, hallucinations, misunderstandings etc. without even fact checking the original articles.

Re: AI assistants misrepresent news content 45% of the time

#17
post #6

Now let's run this experiment against the editorial boards in newsrooms. Obviously, AI isn't an improvement, but people who blindly trust the news have always been credulous rubes. It's just that the alternative is being completely ignorant of the worldviews of everyone around you. Peer-reviewed science is as close as we can get to good consensus and there's a lot of reasons this doesn't work for reporting.

I guess the claim is not that rubes did not used to exist, but rather that technology is increasingly encouraging and streamlining rubism.

I agree with that assessment, or at least that this is indeed the claim.

But, technology also gave us the internet, and social media. Yes, both are used to propagate misinformation, but it also laid bare how bad traditional media was at both a) representing the world competently and b) representing the opinions and views of our neighbors. Manufacturing consent has never been so difficult (or, I suppose, so irrelevant to the actions of the states that claim to represent us).

Re: AI assistants misrepresent news content 45% of the time

#18
post #9

> 45% of all AI answers had at least one significant issue. > 31% of responses showed serious sourcing problems – missing, misleading, or incorrect attributions. > 20% contained major accuracy issues, including hallucinated details and outdated information. I'm generally against whataboutism, but here I think we absolutely have to compare it to human-written news reports. Famously, Michael Crichton introduced the "Ge…

The problem highlighted here is that AI summaries misrepresent the original stories. This just opens a flood gate of slop that is 45% worse than the source, which wasn't stellar to begin with as you point out.

Re: AI assistants misrepresent news content 45% of the time

#19

Now let's run this experiment against the editorial boards in newsrooms. Obviously, AI isn't an improvement, but people who blindly trust the news have always been credulous rubes. It's just that the alternative is being completely ignorant of the worldviews of everyone around you. Peer-reviewed science is as close as we can get to good consensus and there's a lot of reasons this doesn't work for reporting.

> Peer-reviewed science is as close as we can get to good consensus

I think we're on the same side of this, but I just want to say that we can do a lot better. As per studies around the Replication Crisis over the last decade [0], and particularly this 2016 survey conducted by Monya Baker from Nature [1]:

> 1,576 researchers who took a brief online questionnaire on reproducibility found that more than 70% of researchers have tried and failed to reproduce another scientist's experiment results (including 87% of chemists, 77% of biologists, 69% of physicists and engineers, 67% of medical researchers, 64% of earth and environmental scientists, and 62% of all others), and more than half have failed to reproduce their own experiments.

We need to expect better, needing both better incentives and better evaluation, and I think that AI can help with this.

[0] https://en.wikipedia.org/wiki/Replication_crisis

[1] https://www.nature.com/articles/533452a

Re: AI assistants misrepresent news content 45% of the time

#20
I recently tried to get Gemini to collect fresh news and show them to me, and instead of using search it hallucinated everything wholesale, titles, abstracts and links. Not just once, multiple times. I am kind of afraid of using Gemini now for anything related to web search.

Here is a sample:

> [1] Google DeepMind and Harvard researchers propose a new method for testing the ‘theory of mind’ of LLMs - Researchers have introduced a novel framework for evaluating the "theory of mind" capabilities in large language models. Rather than relying on traditional false-belief tasks, this new method assesses an LLM’s ability to infer the mental states of other agents (including other LLMs) within complex social scenarios. It provides a more nuanced benchmark for understanding if these systems are merely mimicking theory of mind through pattern recognition or developing a more robust, generalizable model of other minds. This directly provides material for the construct_metaphysics position by offering a new empirical tool to stress-test the computational foundations of consciousness-related phenomena.

> https://venturebeat.com/ai/google-deepmind-and-harvard-resea...

The link does not work, the title is not found in Google Search either.

Post reply on HN