Live data from Hacker News

AI assistants misrepresent news content 45% of the time

bbc.co.uk

101–110 of 306 posts

Re: AI assistants misrepresent news content 45% of the time

#102
This article should be adjusted to say poor prompting of news content misrepresents news content 45% of the time.

Now, who is responsible for poor prompting?

Maybe the LLM models will just tighten up this part of their models and assistants and suddenly it looks solved.

Re: AI assistants misrepresent news content 45% of the time

#104
post #89

Earlier quoted context omitted.

It did exist but got removed: https://en.wikipedia.org/wiki/Wikipedia:Articles_for_deletio... Quite an omission to not even check for that and it make me think that was done intentionally.

Removed because it was an AI generated article which cited made up sources. Hey, that gives me an idea though, subagents which check whether sources cited exist, and create them whole cloth if they don't

Or subagents that check each link to see if they verify the actual claims the links are sourced for.

Re: AI assistants misrepresent news content 45% of the time

#105

Earlier quoted context omitted.

This is likely because of the knowledge cutoff. I have seen a few cases before of "hallucinations" that turned out to be things that did exist, but no longer do.

The fix for this is for the AI to double-check all links before providing them to the user. I frequently ask ChatGPT to double check that references actually exist when it gives me them. It should be built in!

But that would mean OpenAI would lose even more money on every query.

Re: AI assistants misrepresent news content 45% of the time

#106

I have been unable to recreate any of the failure examples they gave. I don't have co-pilot, but at least Gemini 2.5 pro, ChatGPT5-Thinking, and Perplexity have all give the correct answers as outlined.[1] They don't say what models they were actually using though, so it could be nano models that they asked. They also don't outline the structure of the tests. It seems rigor here was pretty low. Which frankly comes of…

They're talking about assistants, not models, so try using the gemini or perplexity app?

Re: AI assistants misrepresent news content 45% of the time

#107

I am curious if LLMs evangelists understand how off-putting it is when they knee-jerk rationalize how badly these tools are performing. It makes it seem like it isn't about technological capabilities: it is about a religious belief that "competence" is too much to ask of either them or their software tools.

Is that just an LLM thing? I thought that as a society, we decided a long time ago that competence doesn't really matter.

Why else would we be giving high school diplomas to people who can't read at a 5th grade level? Or offshore call center jobs to people who have poor English skills?

Re: AI assistants misrepresent news content 45% of the time

#108
post #42

Earlier quoted context omitted.

> or it (shocking) cites Wikipedia instead of the BBC. No... the problem is that it cites Wikipedia articles that don't exist . > ChatGPT linked to a non-existent Wikipedia article on the “European Union Enlargement Goals for 2040”. In fact, there is no official EU policy under that name. The response hallucinates a URL but also, indirectly, an EU goal and policy.

Actually there was a Wikipedia article of this name, but it was deleted in June -- because it was AI generated. Unfortunately AI falls for this much like humans do. https://en.wikipedia.org/wiki/Wikipedia:Articles_for_deletio...

The biggest problem with that citation isn't that the article has since been deleted. The biggest problem is that that particular Wikipedia article was never a good source in the first place.

That seems to be the real challenge with AI for this use case. It has no real critical thinking skills, so it's not really competent to choose reliable sources. So instead we're lowering the bar to just asking that the sources actually exist. I really hate that. We shouldn't be lowering intellectual standards to meet AI where it's at. These intellectual standards are important and hard-won, and we need to be demanding that AI be the one to rise to meet them.

Re: AI assistants misrepresent news content 45% of the time

#109
post #79
post #20

I recently tried to get Gemini to collect fresh news and show them to me, and instead of using search it hallucinated everything wholesale, titles, abstracts and links. Not just once, multiple times. I am kind of afraid of using Gemini now for anything related to web search. Here is a sample: > [1] Google DeepMind and Harvard researchers propose a new method for testing the ‘theory of mind’ of LLMs - Researchers have…

But LLM can't collect anything. It can generate the most likely characters in a row. What exactly did you expect from it?

Current LLM offerings use realtime web search to collect information and answer questions.

Re: AI assistants misrepresent news content 45% of the time

#110

I am curious if LLMs evangelists understand how off-putting it is when they knee-jerk rationalize how badly these tools are performing. It makes it seem like it isn't about technological capabilities: it is about a religious belief that "competence" is too much to ask of either them or their software tools.

I'm curious if LLM skeptics bother to click through and read the details on a study like this, or if they just reflexively upvote it because it confirms their priors.

This is a hit piece by a media brand that's either feeling threatened or is just incompetent. Or both.

Post reply on HN