Live data from Hacker News

AI assistants misrepresent news content 45% of the time

bbc.co.uk

151–160 of 306 posts

Re: AI assistants misrepresent news content 45% of the time

#151
post #33

Now let's run this experiment against the editorial boards in newsrooms. Obviously, AI isn't an improvement, but people who blindly trust the news have always been credulous rubes. It's just that the alternative is being completely ignorant of the worldviews of everyone around you. Peer-reviewed science is as close as we can get to good consensus and there's a lot of reasons this doesn't work for reporting.

> Now let's run this experiment against the editorial boards in newsrooms. Or against people in general. It's a pet peeve of mine that we get these kinds of articles without a baseline established of how people do on the same measure. Is misrepresenting news content 45% of the time better or worse than the average person? I don't know. By extension: Would a person using an AI assistant misrepresent news more or less…

The difference is the ease with which AI can be rolled out, scaled up, and woven into the fabric of our interactions with society.

Re: AI assistants misrepresent news content 45% of the time

#152
post #128

Earlier quoted context omitted.

I used perplexity for searches and I clicked on all sources that were given. Depending on the model used from 100% to 20% of the urls I tested did not exist. I kept on querying the LLM about it and it finally told me that it generated « the most probable » urls for the topic in question based on the ones he knows exists. Useless.

I share your opinion on the results, but why would you trust the LLM explanation for why it does what it does?

I don’t trust it at all. I wanted to know if he would be able to explain its own results. Just because it was displaying sources and links made me trust it until I checked and was horrified. I wanted to know if it was old link that broke or changed but no apparently

Re: AI assistants misrepresent news content 45% of the time

#153

I am curious if LLMs evangelists understand how off-putting it is when they knee-jerk rationalize how badly these tools are performing. It makes it seem like it isn't about technological capabilities: it is about a religious belief that "competence" is too much to ask of either them or their software tools.

I partially agree, it seems a lot have shifted the argument to news media criticism or something else. But this study is also questionable, for anyone who reads actual academic studies that should be immediately obvious. I don't understand why the bar is this low for some paid Ipsos study vs. some peer-reviewed paper in some IEEE journal?

Like for a study like this I expect as a bare minimum clearly stated model variants used, R@k recall numbers measuring retrieval and something like BLEU or ROUGE to measure summarization accuracy against some baseline on top of their human evaluation metrics. If this is useless for the field itself, I don't understand how this can be useful for anyone outside the field?

Re: AI assistants misrepresent news content 45% of the time

#154

Earlier quoted context omitted.

The biggest problem with that citation isn't that the article has since been deleted. The biggest problem is that that particular Wikipedia article was never a good source in the first place. That seems to be the real challenge with AI for this use case. It has no real critical thinking skills, so it's not really competent to choose reliable sources. So instead we're lowering the bar to just asking that the sources a…

I get what your saying. But you are now asking for a level of intelligence and critical thinking that I honestly believe is higher than the average person. I think its absolutely doable, but I also feel like we shouldn't make it sound like the current behavior is abhorrent or somehow indicative of a failure in the technology.

It's actually great from my point of view - it means we're edging our way into limited superintelligence.

Re: AI assistants misrepresent news content 45% of the time

#155
post #27

According to PEW that's about the same % that trust the BBC's reporting. https://www.pewresearch.org/journalism/fact-sheet/news-media...

You seem to be looking at the wrong chart. Around ~50% each of each politically leaning group use BBC as their primary news source. However, 79% of Brits trust the BBC as per this chart: https://legacy.pewresearch.org/wp-content/uploads/sites/2/20...

That was back in 2017, if I'm reading the chart correctly. A lot has changed since then, so I'd be genuinely curious to see what more recent figures looked like.

Re: AI assistants misrepresent news content 45% of the time

#156

Earlier quoted context omitted.

The fix for this is for the AI to double-check all links before providing them to the user. I frequently ask ChatGPT to double check that references actually exist when it gives me them. It should be built in!

I have found my self doing the same "citation needed" loop - but with ai this is a dangerous game as it will now double down on whatever it made up and go looking for citations to justify its answer. Pre prompting to cite sources is obviously a better way of going about things.

[deleted]

Re: AI assistants misrepresent news content 45% of the time

#157

I'm curious how many people have actually taken the time to compare AI summaries with sources they summarize. I did for a few and ... it was really bad. In my experience, they don't summarize at all, they do a random condensation.. not the same thing at all. In one instance I looked at the result was a key takeaway being the opposite of what it should have been. I don't trust them at all now.

In my experience there is a big difference between good models and weak ones. Quick test with this long article I read recently: https://www.lawfaremedia.org/article/anna--lindsey-halligan-...

The command I ran was `curl -s https://r.jina.ai/https://www.lawfaremedia.org/article/anna-... | cb | ai -m gpt-5-mini summarize this article in one paragraph`. r.jina.ai pulls the text as markdown, and cb just wraps in a ``` code fence, and ai is my own LLM CLI https://github.com/david-crespo/llm-cli.

All of them seem pretty good to me, though at 6 cents the regular use of Sonnet for this purpose would be excessive. Note that reasoning was on the default setting in each case. I think that means the gpt-5 mini one did no reasoning but the other two did.

GPT-5 one paragraph: https://gist.github.com/david-crespo/f2df300ca519c336f9e1953...

GPT-5 three paragraphs: https://gist.github.com/david-crespo/d68f1afaeafdb68771f5103...

GPT-5 mini one paragraph: https://gist.github.com/david-crespo/32512515acc4832f47c3a90...

GPT-5 mini three paragraphs: https://gist.github.com/david-crespo/ed68f09cb70821cffccbf6c...

Sonnet 4.5 one paragraph: https://gist.github.com/david-crespo/e565a82d38699a5bdea4411...

Sonnet 4.5 three paragraphs: https://gist.github.com/david-crespo/2207d8efcc97d754b7d9bf4...

Re: AI assistants misrepresent news content 45% of the time

#160
post #33

Earlier quoted context omitted.

> Now let's run this experiment against the editorial boards in newsrooms. Or against people in general. It's a pet peeve of mine that we get these kinds of articles without a baseline established of how people do on the same measure. Is misrepresenting news content 45% of the time better or worse than the average person? I don't know. By extension: Would a person using an AI assistant misrepresent news more or less…

> It's a pet peeve of mine that we get these kinds of articles without a baseline established of how people do on the same measure I don’t have a personal human news summarizer? The comparison is between a human reading the primary source against the same human reading an LLM hallucination mixed with an LLM referring the primary source. > cynic in me want another question answered too: How often does reporters misrep…

> I don’t have a personal human news summarizer?

Not a personal one. You do however have reporters sitting between you and the source material a lot of the time, and sometimes multiple levels of reporters playing games of telephone with the source material.

> The comparison is between a human reading the primary source against the same human reading an LLM hallucination mixed with an LLM referring the primary source.

In modern news reporting, a fairly substantial proportion of what we digest is not primary sources. It's not at all clear whether an LLM summarising primary sources would be better or worse than reading a reporter passing on primary sources. And in fact, in many cases the news is not even secondary sources - e.g. a wire service report on primary sources getting rewritten by a reporter is not uncommon.

> The fact that you mark as cynical a question answered pretty reliably for most countries sort of tanks the point.

It's a cynical point within the context of this article to point out that it is meaningless to report on the accuracy of AI in isolation because it's not clear that human reporting is better for us. I find it kinda funny that you dismiss this here, after having downplayed the games of telephone that news reporting often is earlier in your reply, thereby making it quite clear I am in fact being a lot more cynical than you about it.

Post reply on HN